Skip to content

Measure what collection narrowing costs before building the classifier - #257

Merged
adamjohnwright merged 2 commits into
mainfrom
009-collection-routing
Sep 19, 2026
Merged

adamjohnwright merged 2 commits into
mainfrom
009-collection-routing

Conversation

@adamjohnwright

@adamjohnwright adamjohnwright commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

The measurement that decides whether spec 009 is worth building, taken now that the machinery is merged and citations actually rank across collections.

Corrected after an adversarial review of this PR — see the second commit. The first measurement said 12/15. It was wrong, and one of the three "failures" was a correct answer. Numbers below are the re-run.

Citation overlap turned out to be arithmetic, not signal

Narrowing to two collections keeps 5–7 of the full top-12 citations; narrowing to one keeps 2. That looks like a large loss — but RRF interleaves the five per-collection lists, so a top-12 draws about 2.4 from each. Keeping two predicts ~4.8, and 5–7 is what I measured.

It measures the interleave, not whether anything useful was lost. I nearly reported it as a finding; it belongs in the record as a warning instead.

Whether answers survive is the measurement that counts

Re-run through run() and report() from answer_sweep itself, with a control arm the first measurement never had. Two needs_live skips on this host, in both arms.

all five collections narrowed to reactions + summations
passed 13 / 13 10 / 13
failed 0 3
skipped 2 2

Two failures are real losses, and they name their collection:

question needs what the narrowed answer said
"What is the UniProt accession for the TP53 protein" (P04637) ewas "not explicitly provided in the context searched"
"Which diseases involve variants of the PTEN gene" disease_variants pathway-level prose, no variant named

The third was not a loss — the check was wrong. With disease_variants excluded, the ABCA1 question named all six curated variants (C1417R, Q537R, S1446L, N935S, W590S, R587W) with the OMIM id. The expectation failed it because the pattern required ABCA1 C1417R adjacency while the answer wrote a numbered list of **C1417R**. The PTEN pattern had the same flaw and would have rejected a correct PTEN answer too. Both are fixed and pinned by a test that fails against the old patterns and still rejects the pathway-level prose the guards were written to catch.

What went wrong the first time

The 12/15 came from a throwaway script that reimplemented the sweep's matching. It never read must, which nine of the fifteen expectations carry, and it counted the two needs_live questions as passes where the sweep skips them. Both errors push the number up.

The re-run records which collections retrieval actually searched and asserts it. That earned itself immediately: the first re-run searched nothing at all — a PoolTimeout on a POSTGRES_LANGGRAPH_DB this host cannot reach — and reported 0/13, which without the check reads as a catastrophic result rather than a void one.

What this settles

Collections are not interchangeable, and the mapping is legible. Accession questions need ewas; PTEN variant questions need disease_variants. That is the signal a classifier can be prompted on.

It is a weaker result than the first measurement suggested, and the weakening is the useful part: content is duplicated across summations more than the guard table assumed. reactions was already known to be unguardable by answer; disease_variants now joins it for the ABCA1 question.

One excluded collection is untested by construction. Narrowing excluded complexes, ewas and disease_variants. No tracked question guards complexes (T005 is open), so its exclusion could not have produced a failure. The absence of a fourth failure is not evidence — and a classifier that never routes to complexes would still score full marks.

It still justifies the fail-wide rule. Where a narrow selection does lose the answer, it removes it rather than degrading it. Widening on uncertainty costs latency; narrowing wrongly costs the answer.

And it sets the acceptance bar: routing ships only if the sweep stays at the control's score with the classifier choosing — 13/13 here, 15/15 in the container where MCP is configured.

Sequencing note

This measurement was only possible because the selection machinery merged first (#255) and the ranking fix (#256) made citations span collections. Taken a day ago it would have shown narrowing as free, because four collections' documents could never be cited anyway.

🤖 Generated with Claude Code

adamjohnwright and others added 2 commits September 19, 2026 02:23
…ections

Two measurements, and only the second means anything.

Citation overlap is arithmetic. Narrowing to two collections keeps 5-7 of the
full top-12, but RRF interleaves five lists so a top-12 draws about 2.4 from
each; keeping two predicts ~4.8. The number measures the interleave, not loss.

Whether answers survive is the measurement that counts. Forcing every tracked
question to reactions+summations: 12 of 15 pass, and the three failures are
exactly the questions whose answers live in the excluded collections -- the
UniProt accession needs ewas, and both variant questions need disease_variants.

That settles three things. Collections are not interchangeable and the mapping is
legible, which is what a classifier can be prompted on. A wrong narrow removes
the answer rather than degrading it -- the failures are missing identifiers, not
vaguer prose -- which is the measured justification for widening on any
uncertainty. And the acceptance bar is 15/15 with the classifier choosing, not
12/15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
EOF
… bad check

The 12/15 in the previous commit was produced by a throwaway script that
reimplemented the sweep's matching. It never read `must`, which nine of the
fifteen expectations carry, and it counted the two `needs_live` questions as
passes where the sweep skips them. Both errors push the number up.

Re-run through `run()` and `report()` from answer_sweep itself, with a control
arm the first measurement never had:

    all five collections     13/13 passed, 2 skipped
    reactions + summations   10/13 passed, 2 skipped, 3 failed

The run now records which collections retrieval actually searched and says so.
That check earned itself immediately: the first re-run searched nothing at all
and reported 0/13, which without it reads as a catastrophic result rather than
a void one.

Two of the three failures are real and name their collection -- TP53's accession
needs `ewas`, PTEN's variants need `disease_variants`. The third was not a loss.
With `disease_variants` excluded, the ABCA1 question named all six curated
variants with the OMIM id, and the expectation failed it because the pattern
required `ABCA1 C1417R` adjacency while the answer wrote a numbered list of
`**C1417R**`. A correct answer, marked as a failure -- and the same pattern in
the PTEN expectation would have rejected a correct answer too.

Both patterns now match a variant without demanding the gene name beside it,
pinned by a test that fails against the old ones and still rejects the
pathway-level prose the guards were written to catch.

So ABCA1's variants are reachable from `summations`, and that question does not
guard `disease_variants`. The guard table says so now. It also records what the
absence of a fourth failure does not mean: `complexes` was excluded too, and no
tracked question guards it, so nothing could have failed for it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 4941dab into main Sep 19, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the 009-collection-routing branch September 19, 2026 03:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant