Skip to the reading
open.com.im the public record, read at the source
Sheet 03 Subject labs.llc/press/ Measured 2026-09-22 22:33–22:42 UTC Method one live search run in a browser, plus an independent read of the upstream

open.com.im › The desk › Sheet 03

Sheet 03 · open data

The Press Room, by the numbers

Every American newspaper the Library of Congress has scanned since 2005, 1789 to 1963, searchable without a key — and then an instrument that asks the archive fifty-one separate questions, one per state, to draw where a word actually ran. We ran it, timed it, and checked its arithmetic against the Library itself.

Chronicling America is the National Digital Newspaper Program: the Library of Congress and the National Endowment for the Humanities paying state institutions, for twenty years, to scan their historic papers to one standard. It is the largest open archive of American newsprint in existence and it sits behind a single keyless JSON API that almost nobody uses directly.

The Press Room uses it directly. No account, no relay, no cached copy of the corpus — every result you see was answered by loc.gov to your browser, and every page image link goes to the Library. That architecture is the review.

State queries

51

fifty states and DC

Gap between starts

1.4 s

about 43 requests a minute

Ink-map floor

71 s

51 × 1.4 s, nothing faster

Live answer

3.9 s

loc.gov, by the page’s own clock

Hits returned

192,049

“airship”, exact to the unit

Printed count is off by

5,082

issues, see the fault below

The benchmark: one search, checked twice

At 22:41:03 UTC we gave the page a word and let it work. It came back with 192,049 page hits, sheet 1 of 9,603, twenty result cards, and its own honest latency stamp: “Read 18:41:00 · loc.gov answered in 3.9 s.” Twenty-seven seconds later we put the same word to the Library’s API ourselves, from the command line, and it answered 192,049.

One live search — the word “airship”, 2026-09-22
MeasureThe pageThe Library, asked directlyRead at (UTC)
Page hits192,049192,04922:41:03 / 22:41:30
Result sheets at 20 per sheet9,6039,603derived, matches
Cards rendered20—22:41:04
Upstream latency3.9 s—page’s own stamp
Our own direct read—1.6 s22:39:46

192,049 ÷ 20 = 9,602.45, so 9,603 sheets. The page’s pagination arithmetic is right, its count is the Library’s count, and it does not round.

The Press Room's results on labs.llc after a live search: a line reading '192,049 page hits · sheet 1 of 9,603' above a grid of nine result cards, each naming a newspaper such as St. Tammany farmer, The Rice belt journal and The Billings gazette, with its town, state and date, an OCR snippet with the word 'airship' highlighted, and a 'Read the page at loc.gov' link.
The result sheet at 22:41:04 UTC on 22 September 2026. Nine of the twenty cards are visible. The OCR is shown exactly as the Library returns it, misreadings and all — “aero nautics”, “nodern Invention”, “alrsps” — which is the correct decision and a rare one.

The instrument

The search is the front door. The ink map is why the page exists. Hand it a word and it asks the archive once per state — fifty states and the District — then shades a hand-drawn map by how many pages carry the word, on a scrubber that re-pools the same counts decade by decade from 1789 to 1963.

Fifty-one requests to a public archive is a rate-limit problem, and the page solves it with a dispatch gate rather than a retry storm: request starts are spaced 1,400 ms apart across every in-flight worker, three at a time, because — as the source says in its own comment — the Library’s refusal is triggered by the burst, not the total.

Figure 1

Seventy-one seconds is the floor, and it is deliberate.

The ink map’s request budget, read from the shipped app.js: 51 states, a 1,400 ms gate between request starts, up to three refusal pauses of 65 s each before it finishes partial.

71.4 s +65 s +65 s +65 s 266.4 s — it stops and says what it missed 51 PACED QUERIES PAUSE 1 PAUSE 2 PAUSE 3 050 s 100 s200 s300 s Three in flight · 1,400 ms between request STARTS · 20 results per sheet CONSTANTS READ FROM press/assets/js/app.js AS SERVED 2026-09-22 22:33 UTC

The page tells visitors to “expect a minute or two”, which is accurate at the floor and optimistic at the ceiling. When a state never answers, its key goes into a failure list and the map is drawn without it rather than with a zero — the distinction most heat maps get wrong.

The Press Room's ink map section on labs.llc, headed 'Where a word ran.' A text field labelled THE WORD TO INK shows the placeholder 'zeppelin, cholera, telegraph…' beside an 'Ink the map' button, above a hand-drawn outline map of the United States with every state unshaded and the status reading 'idle'.
The ink map at 22:38:17 UTC on 22 September 2026, before a word is given to it. The map is drawn from simplified Natural Earth outlines shipped with the page; no tile server is involved.

What the archive cannot tell you, and says so

The best paragraph on the page is the one that undercuts its own map. Coverage is not even — each state chose which papers to scan and how many — so an unshaded state can mean the word never ran there, or that the papers it ran in were never scanned. In the page’s words:

“The map cannot tell those apart, and does not pretend to.” labs.llc/press/, the about band

The same restraint governs the OCR. Snippets are printed exactly as the API returns them, misreadings included, with the Library’s own page scan one click away as the cure. A tool that quietly cleaned up its OCR would read better and lie more.

Where the arithmetic slips

All of which makes the one static number on the page conspicuous. The hero line and the about band both state “3,205,306 issues on the API this page reads”. That figure is a string in the HTML. We read the shipped app.js: nothing in it ever writes to that element.

Figure 2

The only number on the page that is not live is the biggest one.

The collection size printed by labs.llc/press/, against the size the Library’s own API returned to us at 22:39:46 UTC on 2026-09-22. Bar drawn to scale; the gap is drawn to scale too.

THE LIBRARY, 22:39:46 UTC 3,210,388 THE PAGE, AS PRINTED 3,205,306 5,082 issues — 0.16% — one pixel wide at this scale A SMALL ERROR, AND THE ONLY ONE WE FOUND

Both bars start at the same origin and share one scale. The difference is genuinely tiny — that is the point of drawing it honestly rather than exaggerating the axis.

The shortcoming

Two things, and they share a cause. First, the collection count is hard-coded: 3,205,306 in the HTML against 3,210,388 from the Library’s API at 22:39:46 UTC, a gap of 5,082 issues that will only widen, on a page whose whole argument is that it reads the source. The live number costs one request — the same request the page already knows how to make.

Second, and worse: the page cannot tell a rate limit from any other failure, and guesses wrong in one direction. Its own source comment explains why — CORS reduces a 429 to a bare network error — so the code treats every TypeError as rate limiting. We saw this by accident. Our probe blocked the request at the network layer, and at 22:38:52 and again at 22:40:02 UTC the page told us “the Library is rate-limiting this connection — give it a minute”. The Library was doing nothing of the sort; it answered our direct read at 22:39:46 UTC in 1.6 seconds. An offline laptop, a blocked request or a DNS stumble will all be reported to the visitor as the Library’s fault. navigator.onLine and a same-origin liveness ping would separate the two cases, and on a page this careful about blame, the misattribution stands out.

The verdict, such as it is

The Press Room is the most quietly disciplined thing in this trio. It asks the primary archive, it paces itself to that archive’s tolerance rather than its own convenience, it prints what the archive said including the mistakes, and it refuses to fill an unshaded state with a zero it has not earned. Its live count matched the Library to the unit at the moment we checked, which is the only way that sentence ought to be written.

The two faults are both about a number the page does not go and fetch — one it could read, and one it guesses. Neither touches the archive itself. Go and use it: labs.llc/press/. Give the map a word and wait the seventy-one seconds out; it is worth it.