Skip to content

Close the gate hole that answer-level questions cannot close - #263

Merged
adamjohnwright merged 2 commits into
mainfrom
009-t005-complexes-guard
Sep 19, 2026
Merged

adamjohnwright merged 2 commits into
mainfrom
009-t005-complexes-guard

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

This is the T005 work I said should happen before routing goes to beta. It did not end where it was meant to, and the ending is the useful part.

There is no complexes guard question

Six candidates across two question shapes, each asked with the collection present and with it removed from what resolve_collections sees — an exact simulation of a bundle without it, and a lot cheaper than copying 3.3G.

question with without
Which complex contains CYBA and CYBB? 0/4 0/4
Which complex has NUP133, NUP160, NUP37? 4/4 4/4
Which complex is made of LSM10, LSM11, SNRPB? 3/4 4/4

Complex names appear throughout reactions, as the names of inputs and outputs, and throughout summations prose. So any question naming or seeking a complex is answerable without it. The one thing structurally unique to complexes — the component list — cannot be asserted on, because the same configuration returns 1 to 6 of 7 components.

This is the same finding as T007's, reached independently: a collection whose content is duplicated elsewhere cannot be guarded by asking a question, however the question is worded.

Two corrections to my own measurements

A withdrawn claim. A four-run pass appeared to show complexes making the Nup107 answer worse — 1,1,1,1 against 3,3,3,3 without it. It did not survive being run again. The variance between runs is larger than any difference between the arms, so the claim was under-powered and is withdrawn rather than quietly dropped.

A void measurement. My first attempt set the selected_collections ContextVar around the call. That no longer works now the classifier is wired — generate_answer sets it from the classifier's choice and overwrites anything set outside, so both arms searched only complexes. Caught by instrumenting what retrieval actually searched, which is the second time in one day that check has saved a result from being reported.

So the guard moves to where it can exist

One parametrised test over all five real collection names: each is reachable, and selecting it searches nothing else. Closes T005a and T007 together.

It catches what the sweep cannot see by construction — a collection routing can never reach, through a name mismatch or a lookup that silently yields nothing, while every tracked question still passes. Verified by sabotage: making resolve_collections ignore the selection fails all five.

What this means for deploying #262

The reason I gave for holding routing back was that the sweep could not detect a classifier that never routes to complexes. That is now covered at the retrieval level, which is the strongest guard available for that collection. It is a weaker signal than an end-to-end question would have been, and it is the one that exists.

Checks

506 passed, 1 skipped. mypy over all 137 files, ruff, ruff format clean.

🤖 Generated with Claude Code

adamjohnwright and others added 2 commits September 19, 2026 06:00
T005 wanted a question that guards `complexes`. There isn't one, and after
actually trying it the reason is structural rather than a failure of
imagination -- which makes it the same finding as T007's, arrived at twice.

Six candidates across two question shapes, each asked with the collection
present and with it removed from what `resolve_collections` sees, which is an
exact simulation of a bundle without it. Every one answered just as well
without: complex names appear throughout `reactions` as the names of inputs
and outputs, and throughout `summations` prose. The one thing structurally
unique to `complexes` -- the component list -- cannot be asserted on, because
the same configuration returns anywhere from 1 to 6 of 7 components.

An earlier four-run pass appeared to show `complexes` making an answer *worse*,
1,1,1,1 against 3,3,3,3 without it. It did not survive being run again and is
withdrawn. The variance is larger than any difference between the arms.

A false start worth recording: the first attempt set the ContextVar around the
call, which no longer works now the classifier is wired -- `generate_answer`
sets it from the classifier's choice and overwrote it, so both arms searched
only `complexes` and the result was void. Caught by instrumenting what
retrieval actually searched, for the second time in a day.

So the guard moves to where it can exist. One parametrised test over all five
real collection names asserts each is reachable and that selecting it searches
nothing else. That closes T005a and T007 together, and it catches the failure
the sweep cannot see by construction: a collection routing can never reach,
through a name mismatch or a lookup that silently yields nothing, while every
tracked question still passes.

Verified by sabotage -- making `resolve_collections` ignore the selection fails
all five.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…close

Three things, and the third is the one that mattered.

The bundle helper was reconfigured by assigning to a module-level global and
restoring it in `finally`. Nothing breaks today because the suite runs
single-process, but `_bundle` reads that global at call time, so adding
`-n auto` later would let the other test in this file intermittently see a
five-collection bundle and fail for reasons having nothing to do with it. The
collections are a parameter now.

The write-up said "six candidates across two question shapes". It was nine
questions over six complexes -- six of one shape, three of the other. And the
CYBA/CYBB row proves nothing either way, since that question fails in both
arms; it is a bad question, not evidence, and the finding rests on the two
rows that answer. Both corrected.

And the overstatement. I wrote that the retrieval-level assertion covers the
hole I had given as the reason not to deploy routing. It does not. It catches
a *plumbing* failure -- a collection unreachable through a name mismatch or a
lookup that yields nothing. The failure that motivated T005 is a classifier
that never chooses a collection, and nothing deterministic can catch that,
because it is one model call's judgement.

So the residual risk is stated rather than retired: routing can under-serve
`complexes` indefinitely, every tracked question still passes, and the only
symptom is answers that are quietly worse. Smaller than it was, since the
plumbing is pinned and the nine questions found no case where `complexes` was
needed for a correct answer at all. Not nothing, and not what the test checks.
Closing it needs the routing distribution watched over real traffic or a
periodic probe, and neither exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 02166cb into main Sep 19, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the 009-t005-complexes-guard branch September 19, 2026 06:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant