Skip to the reading
open.com.im the public record, read at the source
Sheet 02 Subject labs.llc/patents/ Measured 2026-09-22 22:33–22:46 UTC Method paced public GET, byte counts, shipped pack audit

open.com.im › The desk › Sheet 02

Sheet 02 · open data

The Prior Art, by the numbers

Every classification symbol in the Cooperative Patent Classification — all 254,314 of them — vendored onto one host as 681 shards, so that looking up how a patent office files an invention requires no upstream at all. We weighed the pack and then weighed a single question put to it.

Patent classification is the skeleton key of prior-art searching and it is almost unusably large. The CPC is a hierarchy maintained jointly by the EPO and the USPTO: nine sections at the top, a quarter of a million symbols at the bottom, revised quarterly. Every serious patent search runs through it, and every ordinary search tool asks somebody else’s server what a code means.

The Prior Art does not ask. It ships the scheme.

CPC symbols

254,314

release 2026.08, all of it

Subclass shards

681

one TSV per subclass

Index shards

26

36,677 stems, 843,371 postings

Whole pack

21.7

MiB — 22,728,879 bytes exactly

One lookup

29,643

bytes, median subclass

Keys required

0

verified live, see below

The shape of the scheme

The counts below are not the page’s summary; they are the five levels of meta.json, which the live host served me at 22:35:21 UTC in 18,903 bytes. They sum to 254,314 without remainder, which is the first thing you check when somebody tells you they have vendored a standard.

Figure 1

Nine sections. Two hundred and forty-three thousand leaves.

CPC release 2026.08 by level, drawn on a logarithmic scale because a linear one would render the top four levels invisible. Counted from meta.json, 22:35:21 UTC, 2026-09-22.

SECTIONCLASSSUBCLASS MAIN GROUPSUBGROUP 9137681 9,867243,620 110100 1,00010,000100,000

9 + 137 + 681 + 9,867 + 243,620 = 254,314. The 681 subclasses are the sharding seam: each one becomes its own file, which is why a subclass is also the unit of everything measured on this sheet.

The Prior Art's scheme section on labs.llc, headed 'All of CPC, offline. Words to codes, no model.' A search box reads 'A symbol (H02J, H02J7/00) or words (bicycle brake)' with Look up and All sections buttons. Below, a BROWSE panel lists CPC 2026.08 sections A HUMAN NECESSITIES through G PHYSICS, each with a '+ Sheet' button, beside an empty WORDS to CODES panel.
The scheme browser at 22:37:49 UTC on 22 September 2026. Two panels: codes on the left, words on the right. The right-hand panel is where the byte cost lives, and the page does not say so.

The benchmark: what one question costs

“Offline” is a claim about bytes, so we measured bytes. Opening a subclass requires exactly two files: meta.json once, and that subclass’s TSV. Asking for a word requires meta.json and one index shard, chosen by the first letter of the stem. Those are different costs, and the difference is the most useful thing on this sheet.

Figure 2

A code is cheap. A word is not.

Bytes on the wire for one lookup, each including meta.json at 18,903 B. Linear scale. Shard sizes audited from the shipped pack; the A61B figure was fetched live at 22:35:24 UTC, 2026-09-22.

SUBCLASS, MEDIANWORD SHARD, MEDIAN A61B, MEASURED LIVEWORD SHARD “C”, WORST G05B SUBCLASS, WORST 29,643 128,564 235,246 411,898 596,709 0200,000 400,000600,000 BYTES

Green bars are the codes path; rust bars are the words path. The whole pack is 22,728,879 bytes — 38 times the longest bar here, and 767 times the shortest. That ratio is the entire argument for sharding, and it holds.

So the headline survives contact: you do not download 21.7 MiB to look up a code. The median subclass costs 29,643 bytes, which is a small photograph. The live host answered meta.json in 0.32 s and the 216,343-byte A61B shard in 0.53 s, both first-byte in about 0.24 s.

The pack, audited file by file — 2026-09-22, 22:33–22:35 UTC
ComponentFilesBytesSharePer-file range
Subclass shards (sub/)68117,280,91576.0%70 – 577,806 B, median 10,740
Word index shards (idx/)263,328,03214.6%1,882 – 392,995 B, median 109,661
Skeleton (skeleton.tsv)11,300,6065.7%10,694 rows
Catchwords (catchwords.tsv)1800,4233.5%19,120 rows, 21,731 code refs
Manifest (meta.json)118,9030.1%fetched on every path
Total71022,728,879100%21.68 MiB

Audited against the shipped tree and spot-checked against the live host: four files fetched from labs.llc between 22:35:17 and 22:35:28 UTC matched their shipped byte counts exactly.

What the sharding is actually made of

The index is the part worth admiring. 36,677 stems point at 843,371 postings, and a posting is not a row number — it is a base-36 delta from the previous posting, with an optional span so that a stem found in a parent title is inherited by every child beneath it without being stored again. The stemmer is seven rules long and written out in the manifest in full: ies→y, sses→ss, drop s, drop ing, ied→y, drop ed, drop final e. You can read it, disagree with it, and predict it. No model, no embedding, no vendor.

“Every one of the 254,314 symbols in CPC release 2026.08, served from this site.” labs.llc/patents/, the scheme band

Keyless, checked

The page calls itself keyless, which is the sort of claim that is usually half true because some credential is sitting on a server. There is a relay here — relay.php, for the EPO’s Open Patent Services and the USPTO’s Open Data Portal, both of which do want keys. So I asked it. At 22:45:43 UTC it answered in 84 bytes:

“no EPO OPS token installed — the page reads the keyless record instead” HTTP 501 from labs.llc/patents/relay.php, 2026-09-22 22:45:43 UTC

That is the correct answer. The keyed tier exists, is dormant, and degrades to the public record rather than to an error page. The shipped source says neither token file is ever deployed, and the live host agrees.

The Prior Art's log section on labs.llc, headed 'What was searched, when — and what was not', explaining that every query compiled, every engine opened and every search run is stamped as it happens, and that the list of what the search missed is built from what the log shows was not done.
The log band at 22:38:03 UTC on 22 September 2026. Its “what this search missed” list is assembled only from actions the log shows were not taken — it never claims anything was found.

The shortcoming

The words path is four to fourteen times more expensive than the codes path, and nothing on the page tells you so. Index shards are split by the first letter of the stem, and English technical vocabulary is not evenly distributed across the alphabet: x.tsv is 1,882 bytes while c.tsv is 392,995 — a 209-fold spread. Type a word beginning with c, s or p and you pull between 299 KB and 393 KB before the first suggestion appears; type one beginning with x and you pull two kilobytes. On a slow connection that is the difference between instant and noticeably not. A second-level split, or a small stem dictionary fetched once, would flatten it. The page is precise about everything except its own cost model.

The verdict, such as it is

This is the most ambitious engineering in the estate and it survives being weighed. The scheme really is complete, the shards really do make it cheap, the stemmer really is legible, and the keyless claim really is keyless when you knock on the relay. The honest log — a list built only from what was not done — is the rarest thing here; most search tools are designed to make you feel finished.

Against that: one alphabetic seam in the index, unflattened and unannounced, quietly undoes some of the sharding win on exactly the path most visitors will take first. Fix that and the page has no soft edge left.

Weigh it yourself: labs.llc/patents/. Open the network panel, type a code, then type a word, and watch the difference.