Measure what collection narrowing costs before building the classifier - #257
Merged
Merged
Conversation
…ections Two measurements, and only the second means anything. Citation overlap is arithmetic. Narrowing to two collections keeps 5-7 of the full top-12, but RRF interleaves five lists so a top-12 draws about 2.4 from each; keeping two predicts ~4.8. The number measures the interleave, not loss. Whether answers survive is the measurement that counts. Forcing every tracked question to reactions+summations: 12 of 15 pass, and the three failures are exactly the questions whose answers live in the excluded collections -- the UniProt accession needs ewas, and both variant questions need disease_variants. That settles three things. Collections are not interchangeable and the mapping is legible, which is what a classifier can be prompted on. A wrong narrow removes the answer rather than degrading it -- the failures are missing identifiers, not vaguer prose -- which is the measured justification for widening on any uncertainty. And the acceptance bar is 15/15 with the classifier choosing, not 12/15. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> EOF
… bad check
The 12/15 in the previous commit was produced by a throwaway script that
reimplemented the sweep's matching. It never read `must`, which nine of the
fifteen expectations carry, and it counted the two `needs_live` questions as
passes where the sweep skips them. Both errors push the number up.
Re-run through `run()` and `report()` from answer_sweep itself, with a control
arm the first measurement never had:
all five collections 13/13 passed, 2 skipped
reactions + summations 10/13 passed, 2 skipped, 3 failed
The run now records which collections retrieval actually searched and says so.
That check earned itself immediately: the first re-run searched nothing at all
and reported 0/13, which without it reads as a catastrophic result rather than
a void one.
Two of the three failures are real and name their collection -- TP53's accession
needs `ewas`, PTEN's variants need `disease_variants`. The third was not a loss.
With `disease_variants` excluded, the ABCA1 question named all six curated
variants with the OMIM id, and the expectation failed it because the pattern
required `ABCA1 C1417R` adjacency while the answer wrote a numbered list of
`**C1417R**`. A correct answer, marked as a failure -- and the same pattern in
the PTEN expectation would have rejected a correct answer too.
Both patterns now match a variant without demanding the gene name beside it,
pinned by a test that fails against the old ones and still rejects the
pathway-level prose the guards were written to catch.
So ABCA1's variants are reachable from `summations`, and that question does not
guard `disease_variants`. The guard table says so now. It also records what the
absence of a fourth failure does not mean: `complexes` was excluded too, and no
tracked question guards it, so nothing could have failed for it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The measurement that decides whether spec 009 is worth building, taken now that the machinery is merged and citations actually rank across collections.
Corrected after an adversarial review of this PR — see the second commit. The first measurement said 12/15. It was wrong, and one of the three "failures" was a correct answer. Numbers below are the re-run.
Citation overlap turned out to be arithmetic, not signal
Narrowing to two collections keeps 5–7 of the full top-12 citations; narrowing to one keeps 2. That looks like a large loss — but RRF interleaves the five per-collection lists, so a top-12 draws about 2.4 from each. Keeping two predicts ~4.8, and 5–7 is what I measured.
It measures the interleave, not whether anything useful was lost. I nearly reported it as a finding; it belongs in the record as a warning instead.
Whether answers survive is the measurement that counts
Re-run through
run()andreport()fromanswer_sweepitself, with a control arm the first measurement never had. Twoneeds_liveskips on this host, in both arms.reactions+summationsTwo failures are real losses, and they name their collection:
P04637)ewasdisease_variantsThe third was not a loss — the check was wrong. With
disease_variantsexcluded, the ABCA1 question named all six curated variants (C1417R,Q537R,S1446L,N935S,W590S,R587W) with the OMIM id. The expectation failed it because the pattern requiredABCA1 C1417Radjacency while the answer wrote a numbered list of**C1417R**. The PTEN pattern had the same flaw and would have rejected a correct PTEN answer too. Both are fixed and pinned by a test that fails against the old patterns and still rejects the pathway-level prose the guards were written to catch.What went wrong the first time
The 12/15 came from a throwaway script that reimplemented the sweep's matching. It never read
must, which nine of the fifteen expectations carry, and it counted the twoneeds_livequestions as passes where the sweep skips them. Both errors push the number up.The re-run records which collections retrieval actually searched and asserts it. That earned itself immediately: the first re-run searched nothing at all — a
PoolTimeouton aPOSTGRES_LANGGRAPH_DBthis host cannot reach — and reported 0/13, which without the check reads as a catastrophic result rather than a void one.What this settles
Collections are not interchangeable, and the mapping is legible. Accession questions need
ewas; PTEN variant questions needdisease_variants. That is the signal a classifier can be prompted on.It is a weaker result than the first measurement suggested, and the weakening is the useful part: content is duplicated across
summationsmore than the guard table assumed.reactionswas already known to be unguardable by answer;disease_variantsnow joins it for the ABCA1 question.One excluded collection is untested by construction. Narrowing excluded
complexes,ewasanddisease_variants. No tracked question guardscomplexes(T005 is open), so its exclusion could not have produced a failure. The absence of a fourth failure is not evidence — and a classifier that never routes tocomplexeswould still score full marks.It still justifies the fail-wide rule. Where a narrow selection does lose the answer, it removes it rather than degrading it. Widening on uncertainty costs latency; narrowing wrongly costs the answer.
And it sets the acceptance bar: routing ships only if the sweep stays at the control's score with the classifier choosing — 13/13 here, 15/15 in the container where MCP is configured.
Sequencing note
This measurement was only possible because the selection machinery merged first (#255) and the ranking fix (#256) made citations span collections. Taken a day ago it would have shown narrowing as free, because four collections' documents could never be cited anyway.
🤖 Generated with Claude Code