Repository navigation
Close the gate hole that answer-level questions cannot close - #263
Merged
Merged
Conversation
T005 wanted a question that guards `complexes`. There isn't one, and after actually trying it the reason is structural rather than a failure of imagination -- which makes it the same finding as T007's, arrived at twice. Six candidates across two question shapes, each asked with the collection present and with it removed from what `resolve_collections` sees, which is an exact simulation of a bundle without it. Every one answered just as well without: complex names appear throughout `reactions` as the names of inputs and outputs, and throughout `summations` prose. The one thing structurally unique to `complexes` -- the component list -- cannot be asserted on, because the same configuration returns anywhere from 1 to 6 of 7 components. An earlier four-run pass appeared to show `complexes` making an answer *worse*, 1,1,1,1 against 3,3,3,3 without it. It did not survive being run again and is withdrawn. The variance is larger than any difference between the arms. A false start worth recording: the first attempt set the ContextVar around the call, which no longer works now the classifier is wired -- `generate_answer` sets it from the classifier's choice and overwrote it, so both arms searched only `complexes` and the result was void. Caught by instrumenting what retrieval actually searched, for the second time in a day. So the guard moves to where it can exist. One parametrised test over all five real collection names asserts each is reachable and that selecting it searches nothing else. That closes T005a and T007 together, and it catches the failure the sweep cannot see by construction: a collection routing can never reach, through a name mismatch or a lookup that silently yields nothing, while every tracked question still passes. Verified by sabotage -- making `resolve_collections` ignore the selection fails all five. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…close Three things, and the third is the one that mattered. The bundle helper was reconfigured by assigning to a module-level global and restoring it in `finally`. Nothing breaks today because the suite runs single-process, but `_bundle` reads that global at call time, so adding `-n auto` later would let the other test in this file intermittently see a five-collection bundle and fail for reasons having nothing to do with it. The collections are a parameter now. The write-up said "six candidates across two question shapes". It was nine questions over six complexes -- six of one shape, three of the other. And the CYBA/CYBB row proves nothing either way, since that question fails in both arms; it is a bad question, not evidence, and the finding rests on the two rows that answer. Both corrected. And the overstatement. I wrote that the retrieval-level assertion covers the hole I had given as the reason not to deploy routing. It does not. It catches a *plumbing* failure -- a collection unreachable through a name mismatch or a lookup that yields nothing. The failure that motivated T005 is a classifier that never chooses a collection, and nothing deterministic can catch that, because it is one model call's judgement. So the residual risk is stated rather than retired: routing can under-serve `complexes` indefinitely, every tracked question still passes, and the only symptom is answers that are quietly worse. Smaller than it was, since the plumbing is pinned and the nine questions found no case where `complexes` was needed for a correct answer at all. Not nothing, and not what the test checks. Closing it needs the routing distribution watched over real traffic or a periodic probe, and neither exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is the T005 work I said should happen before routing goes to beta. It did not end where it was meant to, and the ending is the useful part.
There is no
complexesguard questionSix candidates across two question shapes, each asked with the collection present and with it removed from what
resolve_collectionssees — an exact simulation of a bundle without it, and a lot cheaper than copying 3.3G.Complex names appear throughout
reactions, as the names of inputs and outputs, and throughoutsummationsprose. So any question naming or seeking a complex is answerable without it. The one thing structurally unique tocomplexes— the component list — cannot be asserted on, because the same configuration returns 1 to 6 of 7 components.This is the same finding as T007's, reached independently: a collection whose content is duplicated elsewhere cannot be guarded by asking a question, however the question is worded.
Two corrections to my own measurements
A withdrawn claim. A four-run pass appeared to show
complexesmaking the Nup107 answer worse — 1,1,1,1 against 3,3,3,3 without it. It did not survive being run again. The variance between runs is larger than any difference between the arms, so the claim was under-powered and is withdrawn rather than quietly dropped.A void measurement. My first attempt set the
selected_collectionsContextVar around the call. That no longer works now the classifier is wired —generate_answersets it from the classifier's choice and overwrites anything set outside, so both arms searched onlycomplexes. Caught by instrumenting what retrieval actually searched, which is the second time in one day that check has saved a result from being reported.So the guard moves to where it can exist
One parametrised test over all five real collection names: each is reachable, and selecting it searches nothing else. Closes T005a and T007 together.
It catches what the sweep cannot see by construction — a collection routing can never reach, through a name mismatch or a lookup that silently yields nothing, while every tracked question still passes. Verified by sabotage: making
resolve_collectionsignore the selection fails all five.What this means for deploying #262
The reason I gave for holding routing back was that the sweep could not detect a classifier that never routes to
complexes. That is now covered at the retrieval level, which is the strongest guard available for that collection. It is a weaker signal than an end-to-end question would have been, and it is the one that exists.Checks
506 passed, 1 skipped. mypy over all 137 files, ruff, ruff format clean.
🤖 Generated with Claude Code