open.com.im › The desk › Sheet 02
Sheet 02 · open data
The Prior Art, by the numbers
Every classification symbol in the Cooperative Patent Classification — all 254,314 of them — vendored onto one host as 681 shards, so that looking up how a patent office files an invention requires no upstream at all. We weighed the pack and then weighed a single question put to it.
Patent classification is the skeleton key of prior-art searching and it is almost unusably large. The CPC is a hierarchy maintained jointly by the EPO and the USPTO: nine sections at the top, a quarter of a million symbols at the bottom, revised quarterly. Every serious patent search runs through it, and every ordinary search tool asks somebody else’s server what a code means.
The Prior Art does not ask. It ships the scheme.
CPC symbols
254,314
release 2026.08, all of it
Subclass shards
681
one TSV per subclass
Index shards
26
36,677 stems, 843,371 postings
Whole pack
21.7
MiB — 22,728,879 bytes exactly
One lookup
29,643
bytes, median subclass
Keys required
0
verified live, see below
The shape of the scheme
The counts below are not the page’s summary; they are the five levels of
meta.json, which the live host served me at 22:35:21 UTC in 18,903 bytes. They
sum to 254,314 without remainder, which is the first thing you check when somebody tells you
they have vendored a standard.
Figure 1
Nine sections. Two hundred and forty-three thousand leaves.
CPC release 2026.08 by level, drawn on a logarithmic scale because a linear one would render the top four levels invisible. Counted from meta.json, 22:35:21 UTC, 2026-09-22.
9 + 137 + 681 + 9,867 + 243,620 = 254,314. The 681 subclasses are the sharding seam: each one becomes its own file, which is why a subclass is also the unit of everything measured on this sheet.
The benchmark: what one question costs
“Offline” is a claim about bytes, so we measured bytes. Opening a subclass
requires exactly two files: meta.json once, and that subclass’s TSV. Asking
for a word requires meta.json and one index shard, chosen by the first letter of
the stem. Those are different costs, and the difference is the most useful thing on this sheet.
Figure 2
A code is cheap. A word is not.
Bytes on the wire for one lookup, each including meta.json at 18,903 B. Linear scale. Shard sizes audited from the shipped pack; the A61B figure was fetched live at 22:35:24 UTC, 2026-09-22.
Green bars are the codes path; rust bars are the words path. The whole pack is 22,728,879 bytes — 38 times the longest bar here, and 767 times the shortest. That ratio is the entire argument for sharding, and it holds.
So the headline survives contact: you do not download 21.7 MiB to look up a code. The
median subclass costs 29,643 bytes, which is a small photograph. The live host answered
meta.json in 0.32 s and the 216,343-byte A61B shard in 0.53 s, both first-byte in
about 0.24 s.
| Component | Files | Bytes | Share | Per-file range |
|---|---|---|---|---|
Subclass shards (sub/) | 681 | 17,280,915 | 76.0% | 70 – 577,806 B, median 10,740 |
Word index shards (idx/) | 26 | 3,328,032 | 14.6% | 1,882 – 392,995 B, median 109,661 |
Skeleton (skeleton.tsv) | 1 | 1,300,606 | 5.7% | 10,694 rows |
Catchwords (catchwords.tsv) | 1 | 800,423 | 3.5% | 19,120 rows, 21,731 code refs |
Manifest (meta.json) | 1 | 18,903 | 0.1% | fetched on every path |
| Total | 710 | 22,728,879 | 100% | 21.68 MiB |
Audited against the shipped tree and spot-checked against the live host: four files fetched from labs.llc between 22:35:17 and 22:35:28 UTC matched their shipped byte counts exactly.
What the sharding is actually made of
The index is the part worth admiring. 36,677 stems point at 843,371 postings, and a posting is not a row number — it is a base-36 delta from the previous posting, with an optional span so that a stem found in a parent title is inherited by every child beneath it without being stored again. The stemmer is seven rules long and written out in the manifest in full: ies→y, sses→ss, drop s, drop ing, ied→y, drop ed, drop final e. You can read it, disagree with it, and predict it. No model, no embedding, no vendor.
Keyless, checked
The page calls itself keyless, which is the sort of claim that is usually half true because
some credential is sitting on a server. There is a relay here — relay.php,
for the EPO’s Open Patent Services and the USPTO’s Open Data Portal, both of which
do want keys. So I asked it. At 22:45:43 UTC it answered in 84 bytes:
That is the correct answer. The keyed tier exists, is dormant, and degrades to the public record rather than to an error page. The shipped source says neither token file is ever deployed, and the live host agrees.
The shortcoming
The words path is four to fourteen times more expensive than the codes path, and nothing
on the page tells you so. Index shards are split by the first letter of the stem, and English
technical vocabulary is not evenly distributed across the alphabet: x.tsv is
1,882 bytes while c.tsv is 392,995 — a 209-fold spread. Type a word
beginning with c, s or p and you pull between 299 KB and 393 KB
before the first suggestion appears; type one beginning with x and you pull two
kilobytes. On a slow connection that is the difference between instant and noticeably
not. A second-level split, or a small stem dictionary fetched once, would flatten it. The
page is precise about everything except its own cost model.
The verdict, such as it is
This is the most ambitious engineering in the estate and it survives being weighed. The scheme really is complete, the shards really do make it cheap, the stemmer really is legible, and the keyless claim really is keyless when you knock on the relay. The honest log — a list built only from what was not done — is the rarest thing here; most search tools are designed to make you feel finished.
Against that: one alphabetic seam in the index, unflattened and unannounced, quietly undoes some of the sharding win on exactly the path most visitors will take first. Fix that and the page has no soft edge left.
Weigh it yourself: labs.llc/patents/. Open the network panel, type a code, then type a word, and watch the difference.