Repository navigation
Watch the routing decision, because no answer can - #267
Merged
Merged
Conversation
T005b wanted the observation no test can make. `answer-sweep` cannot see a classifier that stops choosing a collection: neither `complexes` nor `reactions` can be guarded by asking a question, since their content is duplicated in `summations` prose and in the input/output names `reactions` carries. A collection that goes unchosen produces answers that are confident, plausible and slightly worse, and every tracked question still passes. `bin/routing-probe` asks the classifier ten questions, two per collection, and checks what it *selected* rather than what any answer said. One model call each, cheap enough to run on every deploy. Two properties, with the same asymmetry the feature rests on. Coverage: every collection chosen by at least one question -- the guard for `complexes`, which nothing else watches. No wrong narrow: a question that narrows must include the collection its answer lives in. An empty selection is never a failure, because widening is safe and the prompt asks for it when the model is unsure. First run: 10/10 correct, every collection chosen, nothing left open. So `complexes` is being routed to and the risk that argued against deploying is not currently realised. The tests check that the guard fires rather than that the happy path is quiet, and the coverage assertion was verified by sabotage -- forcing `missing` empty fails the test that exists to catch it. Two probes per collection is itself asserted, so one passing is not luck, and so is the existence of a probe per collection: a collection with no probe cannot be reported as never chosen, which would make this file quietly useless for exactly the collection it was written for. T005c is open and matters more than this commit: one run is a snapshot of one model's judgement on one day. The risk was always drift, and the probe is only a guard if it keeps being run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This is T005b — the guard I said should exist before collection routing goes to beta, and the reason I kept advising against deploying it.
The gap
answer-sweepcannot see a classifier that stops choosing a collection. Neithercomplexesnorreactionscan be guarded by asking a question — measured — because their content is duplicated insummationsprose and in the input/output namesreactionscarries. So a collection that goes unchosen yields answers that are confident, plausible and slightly worse, while every tracked question passes.The probe
bin/routing-probeasks the classifier ten questions, two per collection, and checks what it selected rather than what any answer said. One model call each — cheap enough to run on every deploy.complexes, which nothing else watches.First run
complexesis being routed to. The under-routing risk that argued against deploying is not currently realised.What the tests check
That the guard fires, not that the happy path is quiet — verified by sabotage: forcing
missingempty fails the test written to catch it. Two probes per collection is itself asserted so one passing is not luck, and so is the existence of a probe per collection, since a collection with no probe cannot be reported as never chosen — which would make this file quietly useless for exactly the collection it was written for.T005c is open and matters more than this PR
One run is a snapshot of one model's judgement on one day. The risk was always drift: a classifier that stops choosing
complexesnext month gives you confident answers and a green sweep. The probe is only a guard if it keeps being run, which is why the follow-up is to run it on every deploy beside the sweep.513 passed, 1 skipped; mypy over 139 files, ruff clean.
🤖 Generated with Claude Code