Skip to content

Watch the routing decision, because no answer can - #267

Merged
adamjohnwright merged 1 commit into
mainfrom
009-routing-probe
Sep 19, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
009-routing-probe

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

This is T005b — the guard I said should exist before collection routing goes to beta, and the reason I kept advising against deploying it.

The gap

answer-sweep cannot see a classifier that stops choosing a collection. Neither complexes nor reactions can be guarded by asking a question — measured — because their content is duplicated in summations prose and in the input/output names reactions carries. So a collection that goes unchosen yields answers that are confident, plausible and slightly worse, while every tracked question passes.

The probe

bin/routing-probe asks the classifier ten questions, two per collection, and checks what it selected rather than what any answer said. One model call each — cheap enough to run on every deploy.

  • Coverage: every collection chosen by at least one question. The guard for complexes, which nothing else watches.
  • No wrong narrow: a question that narrows must include the collection its answer lives in.
  • An empty selection is never a failure. Widening is the safe direction and the prompt asks for it when the model is unsure. Open selections are counted and reported, not failed.

First run

10 narrowed, 0 left open, 0 narrowed wrongly, 0 errored
every collection was chosen by at least one question

complexes is being routed to. The under-routing risk that argued against deploying is not currently realised.

What the tests check

That the guard fires, not that the happy path is quiet — verified by sabotage: forcing missing empty fails the test written to catch it. Two probes per collection is itself asserted so one passing is not luck, and so is the existence of a probe per collection, since a collection with no probe cannot be reported as never chosen — which would make this file quietly useless for exactly the collection it was written for.

T005c is open and matters more than this PR

One run is a snapshot of one model's judgement on one day. The risk was always drift: a classifier that stops choosing complexes next month gives you confident answers and a green sweep. The probe is only a guard if it keeps being run, which is why the follow-up is to run it on every deploy beside the sweep.

513 passed, 1 skipped; mypy over 139 files, ruff clean.

🤖 Generated with Claude Code

T005b wanted the observation no test can make. `answer-sweep` cannot see a
classifier that stops choosing a collection: neither `complexes` nor
`reactions` can be guarded by asking a question, since their content is
duplicated in `summations` prose and in the input/output names `reactions`
carries. A collection that goes unchosen produces answers that are confident,
plausible and slightly worse, and every tracked question still passes.

`bin/routing-probe` asks the classifier ten questions, two per collection, and
checks what it *selected* rather than what any answer said. One model call
each, cheap enough to run on every deploy.

Two properties, with the same asymmetry the feature rests on. Coverage: every
collection chosen by at least one question -- the guard for `complexes`, which
nothing else watches. No wrong narrow: a question that narrows must include
the collection its answer lives in. An empty selection is never a failure,
because widening is safe and the prompt asks for it when the model is unsure.

First run: 10/10 correct, every collection chosen, nothing left open. So
`complexes` is being routed to and the risk that argued against deploying is
not currently realised.

The tests check that the guard fires rather than that the happy path is quiet,
and the coverage assertion was verified by sabotage -- forcing `missing` empty
fails the test that exists to catch it. Two probes per collection is itself
asserted, so one passing is not luck, and so is the existence of a probe per
collection: a collection with no probe cannot be reported as never chosen,
which would make this file quietly useless for exactly the collection it was
written for.

T005c is open and matters more than this commit: one run is a snapshot of one
model's judgement on one day. The risk was always drift, and the probe is only
a guard if it keeps being run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit cea10c3 into main Sep 19, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the 009-routing-probe branch September 19, 2026 12:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant