Rewritten 2026-09-27 after the linking reassessment. This issue is the measurement instrument for sense-level shelf coverage and match quality. It fetches nothing for the whole dictionary and runs no ladder pass of its own; the coverage work moved to #217 (see Relationship). The earlier third box, "run the ladder over at least one new scope", is withdrawn: the identifier rungs are exhausted against today's registry (328 lexemes to gain), and the gap is the 21,544 entities the sources name that were never minted.
Purpose
#172 shipped the mechanism: the concept hop, the word-level tier, promotion and the honest empty. It did not establish its title's claim that Quotes, Artworks and Images "appear on most pages". This issue says, with a denominator and a number, how many pages are eligible for a sense-level shelf, how many retrieve, how many retrieve something relevant, and how precise the two inferred tiers are. Every claim about coverage on this project is allowed only with this issue's sample beside it.
Where it stands (devils_dictionary_v2, 2026-09-27, after the #188 restore and promotion runs 206–211)
| English |
lexemes |
pages (lemmas) |
| total |
1,541,669 |
1,407,971 |
| with an active sense-backed Wikidata link |
20,700 (1.34 %) |
20,246 (1.44 %) |
… only through promotion (corroborated_gloss) |
5,690 |
5,570 |
| word-level only (candidate ≥ 0.85, no sense link) |
— |
2,571 (0.18 %) |
| naming a source-published QID absent from the registry |
40,370 |
— |
These are pages eligible for a sense-level shelf. Retrieval is measured only on samples. Of the 13 probe words, 5 show a Quotes shelf. Candidates exist only where the ladder has run (animals, emotions, culture); none of 50 random common nouns had one. Detail: docs/integrations/wikiquote.md, Final sweep; scripts and SQL: docs/spikes/2026-09-27-issue-172-sweep/ (probes.sql, run_pages.exs, baseline.sql are the starting point here).
The instrument
Denominator. Choose and commit one: all English pages, or a fixed list of common words (for example the top 10,000 lemmas by a named frequency source). Commit the list or its query under docs/spikes/.
Sample. A fixed random sample of 200 pages from that denominator, seeded and committed, re-run unchanged after every change listed under Checkpoints. Report, per page and in total:
| column |
meaning |
| eligible |
has a tier-1 sense link or a tier-2 candidate ≥ 0.85 |
| retrieves |
each of Quotes / Artworks / Images returns ≥ 1 result when run in-process (run_pages.exs) |
| level |
sense or word |
| minor-sense |
the shelf follows a sense that is not the page's first (the power → business magnate case) |
| relevant |
a person's judgment on the first three results, with the rubric written down |
The sample is an evaluation. It is not an instruction to fetch content for the whole denominator, and enlarging it is not a serving strategy: pages are served by the visit-driven lifecycle in Discovery, which runs on the words people open.
Precision reviews, each with a written rubric and a 100-row random sample, committed as CSV:
Checkpoints
Relationship
Purpose
#172 shipped the mechanism: the concept hop, the word-level tier, promotion and the honest empty. It did not establish its title's claim that Quotes, Artworks and Images "appear on most pages". This issue says, with a denominator and a number, how many pages are eligible for a sense-level shelf, how many retrieve, how many retrieve something relevant, and how precise the two inferred tiers are. Every claim about coverage on this project is allowed only with this issue's sample beside it.
Where it stands (
devils_dictionary_v2, 2026-09-27, after the #188 restore and promotion runs 206–211)corroborated_gloss)These are pages eligible for a sense-level shelf. Retrieval is measured only on samples. Of the 13 probe words, 5 show a Quotes shelf. Candidates exist only where the ladder has run (animals, emotions, culture); none of 50 random common nouns had one. Detail:
docs/integrations/wikiquote.md, Final sweep; scripts and SQL:docs/spikes/2026-09-27-issue-172-sweep/(probes.sql,run_pages.exs,baseline.sqlare the starting point here).The instrument
Denominator. Choose and commit one: all English pages, or a fixed list of common words (for example the top 10,000 lemmas by a named frequency source). Commit the list or its query under
docs/spikes/.Sample. A fixed random sample of 200 pages from that denominator, seeded and committed, re-run unchanged after every change listed under Checkpoints. Report, per page and in total:
run_pages.exs)The sample is an evaluation. It is not an instruction to fetch content for the whole denominator, and enlarging it is not a serving strategy: pages are served by the visit-driven lifecycle in
Discovery, which runs on the words people open.Precision reviews, each with a written rubric and a 100-row random sample, committed as CSV:
corroborated_glosslinks. Target ≥ 90 %. Below it, promotion's rule changes or Discovery relation assessor: lite ML and LLM judging of provider results before persona curation #101 takes it over (Sense-level shelves show on 1 % of words: the concept hop, the word-level tier from corroborated candidates, and promotion — so Quotes, Artworks and Images appear on most pages without a name match #172 decision 3 said this would decide).Checkpoints
Relationship