open.com.im › The desk › Sheet 03
Sheet 03 · open data
The Press Room, by the numbers
Every American newspaper the Library of Congress has scanned since 2005, 1789 to 1963, searchable without a key — and then an instrument that asks the archive fifty-one separate questions, one per state, to draw where a word actually ran. We ran it, timed it, and checked its arithmetic against the Library itself.
Chronicling America is the National Digital Newspaper Program: the Library of Congress and the National Endowment for the Humanities paying state institutions, for twenty years, to scan their historic papers to one standard. It is the largest open archive of American newsprint in existence and it sits behind a single keyless JSON API that almost nobody uses directly.
The Press Room uses it directly. No account, no relay, no cached copy of the corpus — every result you see was answered by loc.gov to your browser, and every page image link goes to the Library. That architecture is the review.
State queries
51
fifty states and DC
Gap between starts
1.4 s
about 43 requests a minute
Ink-map floor
71 s
51 × 1.4 s, nothing faster
Live answer
3.9 s
loc.gov, by the page’s own clock
Hits returned
192,049
“airship”, exact to the unit
Printed count is off by
5,082
issues, see the fault below
The benchmark: one search, checked twice
At 22:41:03 UTC we gave the page a word and let it work. It came back with 192,049 page hits, sheet 1 of 9,603, twenty result cards, and its own honest latency stamp: “Read 18:41:00 · loc.gov answered in 3.9 s.” Twenty-seven seconds later we put the same word to the Library’s API ourselves, from the command line, and it answered 192,049.
| Measure | The page | The Library, asked directly | Read at (UTC) |
|---|---|---|---|
| Page hits | 192,049 | 192,049 | 22:41:03 / 22:41:30 |
| Result sheets at 20 per sheet | 9,603 | 9,603 | derived, matches |
| Cards rendered | 20 | — | 22:41:04 |
| Upstream latency | 3.9 s | — | page’s own stamp |
| Our own direct read | — | 1.6 s | 22:39:46 |
192,049 ÷ 20 = 9,602.45, so 9,603 sheets. The page’s pagination arithmetic is right, its count is the Library’s count, and it does not round.
The instrument
The search is the front door. The ink map is why the page exists. Hand it a word and it asks the archive once per state — fifty states and the District — then shades a hand-drawn map by how many pages carry the word, on a scrubber that re-pools the same counts decade by decade from 1789 to 1963.
Fifty-one requests to a public archive is a rate-limit problem, and the page solves it with a dispatch gate rather than a retry storm: request starts are spaced 1,400 ms apart across every in-flight worker, three at a time, because — as the source says in its own comment — the Library’s refusal is triggered by the burst, not the total.
Figure 1
Seventy-one seconds is the floor, and it is deliberate.
The ink map’s request budget, read from the shipped app.js: 51 states, a 1,400 ms gate between request starts, up to three refusal pauses of 65 s each before it finishes partial.
The page tells visitors to “expect a minute or two”, which is accurate at the floor and optimistic at the ceiling. When a state never answers, its key goes into a failure list and the map is drawn without it rather than with a zero — the distinction most heat maps get wrong.
What the archive cannot tell you, and says so
The best paragraph on the page is the one that undercuts its own map. Coverage is not even — each state chose which papers to scan and how many — so an unshaded state can mean the word never ran there, or that the papers it ran in were never scanned. In the page’s words:
The same restraint governs the OCR. Snippets are printed exactly as the API returns them, misreadings included, with the Library’s own page scan one click away as the cure. A tool that quietly cleaned up its OCR would read better and lie more.
Where the arithmetic slips
All of which makes the one static number on the page conspicuous. The hero line and the
about band both state “3,205,306 issues on the API this page reads”. That figure is
a string in the HTML. We read the shipped app.js: nothing in it ever writes to
that element.
Figure 2
The only number on the page that is not live is the biggest one.
The collection size printed by labs.llc/press/, against the size the Library’s own API returned to us at 22:39:46 UTC on 2026-09-22. Bar drawn to scale; the gap is drawn to scale too.
Both bars start at the same origin and share one scale. The difference is genuinely tiny — that is the point of drawing it honestly rather than exaggerating the axis.
The shortcoming
Two things, and they share a cause. First, the collection count is hard-coded: 3,205,306 in the HTML against 3,210,388 from the Library’s API at 22:39:46 UTC, a gap of 5,082 issues that will only widen, on a page whose whole argument is that it reads the source. The live number costs one request — the same request the page already knows how to make.
Second, and worse: the page cannot tell a rate limit from any other failure, and guesses
wrong in one direction. Its own source comment explains why — CORS reduces a 429 to a
bare network error — so the code treats every TypeError as rate limiting. We
saw this by accident. Our probe blocked the request at the network layer, and at 22:38:52
and again at 22:40:02 UTC the page told us “the Library is rate-limiting this
connection — give it a minute”. The Library was doing nothing of the sort; it
answered our direct read at 22:39:46 UTC in 1.6 seconds. An offline laptop, a blocked
request or a DNS stumble will all be reported to the visitor as the Library’s fault.
navigator.onLine and a same-origin liveness ping would separate the two cases,
and on a page this careful about blame, the misattribution stands out.
The verdict, such as it is
The Press Room is the most quietly disciplined thing in this trio. It asks the primary archive, it paces itself to that archive’s tolerance rather than its own convenience, it prints what the archive said including the mistakes, and it refuses to fill an unshaded state with a zero it has not earned. Its live count matched the Library to the unit at the moment we checked, which is the only way that sentence ought to be written.
The two faults are both about a number the page does not go and fetch — one it could read, and one it guesses. Neither touches the archive itself. Go and use it: labs.llc/press/. Give the map a word and wait the seventy-one seconds out; it is worth it.