Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions specs/009-collection-routing/research.md
Original file line number Diff line number Diff line change
Expand Up @@ -399,3 +399,15 @@ deploying is not currently realised.
What that is not: a guarantee. It is a snapshot of one model's judgement on
one day. The risk was always drift, and the probe is only a guard if it keeps
being run -- which is why T005c exists and is open.

The probe runs on every deploy from 2026-09-19, wired into
`~/update-beta-chat.sh` beside the sweep. That is what turns it from a
snapshot into a guard: the risk was always drift, and a check run once is a
measurement rather than a control.

It fails the deploy and points at `--rollback` rather than rolling back by
itself, which is the judgement the sweep already makes -- a rollback is
disruptive too, the container is already serving, and which is worse depends
on the failure. An image built before the probe existed skips it with a
warning instead of failing, so the check cannot block a rollback to an older
tag.
2 changes: 1 addition & 1 deletion specs/009-collection-routing/tasks.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ acceptance criterion cannot detect the failure the feature can cause.
- [x] T005 [P] ~~Find a `complexes` candidate that fails without the collection~~ **Attempted and abandoned with a reason, 2026-09-19.** Nine questions over six complexes, in two shapes; every one answered just as well without the collection, because complex names appear throughout `reactions` (as input and output names) and `summations` prose. The one thing unique to `complexes` -- the component list -- has answers too variable to assert on: the same configuration returned 1 to 6 of 7 components. See research.md
- [x] T005a [P] Assert at retrieval level that `complexes` was searched, in tests/retrievers/test_sync_async_equivalence.py -- parametrised over all five real collection names: each is reachable, and selecting it searches nothing else. Verified by sabotaging `resolve_collections` to ignore the selection, which fails all five. **Catches a plumbing failure, not a routing one**: a classifier that never chooses a collection is still undetectable, and needs a production probe rather than a test (research.md)
- [x] T005b Probe the routing decision, since no answer can: `bin/routing-probe` asks the classifier ten questions, two per collection, and checks **what was selected** rather than what any answer said. Fails if a collection is never chosen, or if a question narrows away from where its answer lives; an open selection is never a failure, because widening is the safe direction. First run 2026-09-19: 10/10 correct, every collection chosen
- [ ] T005c Run the probe on every deploy, beside the sweep. **The residual risk is not that the probe is missing, it is that under-routing drifts in between runs** -- a classifier that stops choosing `complexes` next month produces confident plausible answers and a green sweep. One run is a snapshot; the guard is the repetition
- [x] T005c Run the probe on every deploy, beside the sweep. Wired into `~/update-beta-chat.sh` on 2026-09-19, immediately after the sweep and honouring the same `--skip-sweep` flag. Fails the deploy and points at `--rollback`, deliberately not rolling back automatically -- the same judgement the sweep makes, since a rollback is disruptive and the container is already serving. Skips with a warning on an image built before the probe existed, which is what the currently deployed image does. **That script is not version controlled**, so this task is its only record in the repository
- [x] T006 [P] Add an `ewas`-dependent question to `EXPECTATIONS` in src/evaluation/answer_sweep.py (TP53 UniProt P04637; verified by removal)
- [x] T007 Assert at retrieval level that `reactions` was searched -- closed by the same parametrised test as T005a, in tests/retrievers/test_sync_async_equivalence.py. No answer-level question can guard it, because every reaction name also appears in `summations`
- [x] T008 Verify each new question FAILS when its collection is removed from the bundle copy, and passes with it present; record the evidence in the PR (method established; two of four candidates survived it)
Expand Down
Loading