diff --git a/specs/009-collection-routing/research.md b/specs/009-collection-routing/research.md index fcee6ce..fdc1f3c 100644 --- a/specs/009-collection-routing/research.md +++ b/specs/009-collection-routing/research.md @@ -399,3 +399,15 @@ deploying is not currently realised. What that is not: a guarantee. It is a snapshot of one model's judgement on one day. The risk was always drift, and the probe is only a guard if it keeps being run -- which is why T005c exists and is open. + +The probe runs on every deploy from 2026-09-19, wired into +`~/update-beta-chat.sh` beside the sweep. That is what turns it from a +snapshot into a guard: the risk was always drift, and a check run once is a +measurement rather than a control. + +It fails the deploy and points at `--rollback` rather than rolling back by +itself, which is the judgement the sweep already makes -- a rollback is +disruptive too, the container is already serving, and which is worse depends +on the failure. An image built before the probe existed skips it with a +warning instead of failing, so the check cannot block a rollback to an older +tag. diff --git a/specs/009-collection-routing/tasks.md b/specs/009-collection-routing/tasks.md index 1448182..f7c5173 100644 --- a/specs/009-collection-routing/tasks.md +++ b/specs/009-collection-routing/tasks.md @@ -22,7 +22,7 @@ acceptance criterion cannot detect the failure the feature can cause. - [x] T005 [P] ~~Find a `complexes` candidate that fails without the collection~~ **Attempted and abandoned with a reason, 2026-09-19.** Nine questions over six complexes, in two shapes; every one answered just as well without the collection, because complex names appear throughout `reactions` (as input and output names) and `summations` prose. The one thing unique to `complexes` -- the component list -- has answers too variable to assert on: the same configuration returned 1 to 6 of 7 components. See research.md - [x] T005a [P] Assert at retrieval level that `complexes` was searched, in tests/retrievers/test_sync_async_equivalence.py -- parametrised over all five real collection names: each is reachable, and selecting it searches nothing else. Verified by sabotaging `resolve_collections` to ignore the selection, which fails all five. **Catches a plumbing failure, not a routing one**: a classifier that never chooses a collection is still undetectable, and needs a production probe rather than a test (research.md) - [x] T005b Probe the routing decision, since no answer can: `bin/routing-probe` asks the classifier ten questions, two per collection, and checks **what was selected** rather than what any answer said. Fails if a collection is never chosen, or if a question narrows away from where its answer lives; an open selection is never a failure, because widening is the safe direction. First run 2026-09-19: 10/10 correct, every collection chosen -- [ ] T005c Run the probe on every deploy, beside the sweep. **The residual risk is not that the probe is missing, it is that under-routing drifts in between runs** -- a classifier that stops choosing `complexes` next month produces confident plausible answers and a green sweep. One run is a snapshot; the guard is the repetition +- [x] T005c Run the probe on every deploy, beside the sweep. Wired into `~/update-beta-chat.sh` on 2026-09-19, immediately after the sweep and honouring the same `--skip-sweep` flag. Fails the deploy and points at `--rollback`, deliberately not rolling back automatically -- the same judgement the sweep makes, since a rollback is disruptive and the container is already serving. Skips with a warning on an image built before the probe existed, which is what the currently deployed image does. **That script is not version controlled**, so this task is its only record in the repository - [x] T006 [P] Add an `ewas`-dependent question to `EXPECTATIONS` in src/evaluation/answer_sweep.py (TP53 UniProt P04637; verified by removal) - [x] T007 Assert at retrieval level that `reactions` was searched -- closed by the same parametrised test as T005a, in tests/retrievers/test_sync_async_equivalence.py. No answer-level question can guard it, because every reaction name also appears in `summations` - [x] T008 Verify each new question FAILS when its collection is removed from the bundle copy, and passes with it present; record the evidence in the PR (method established; two of four candidates survived it)