From 855d968bfb4ac59f6f9eefe3de2802bf2b66ab99 Mon Sep 17 00:00:00 2001 From: Adam Wright Date: Sat, 19 Sep 2026 06:26:54 +0000 Subject: [PATCH] Leave the risk T005 was written for open, because it is T005a is checked and the classifier risk is not handled, which is the state most likely to be misread. A register that says "handled" is why nobody looks again -- and a classifier that never selects a collection produces confident plausible answers rather than an error anyone notices. T005b records it as open with what closing it would take: the routing distribution observed over real traffic, or a periodic probe. Neither is a test, which is why it cannot be checked off by adding one. Co-Authored-By: Claude Opus 5 --- specs/009-collection-routing/tasks.md | 1 + 1 file changed, 1 insertion(+) diff --git a/specs/009-collection-routing/tasks.md b/specs/009-collection-routing/tasks.md index c502b1a..e4f2766 100644 --- a/specs/009-collection-routing/tasks.md +++ b/specs/009-collection-routing/tasks.md @@ -21,6 +21,7 @@ acceptance criterion cannot detect the failure the feature can cause. - [x] T004 [P] Add a `summations`-dependent question to `EXPECTATIONS` in src/evaluation/answer_sweep.py (Selective autophagy / lysosome; verified by removal) - [x] T005 [P] ~~Find a `complexes` candidate that fails without the collection~~ **Attempted and abandoned with a reason, 2026-09-19.** Nine questions over six complexes, in two shapes; every one answered just as well without the collection, because complex names appear throughout `reactions` (as input and output names) and `summations` prose. The one thing unique to `complexes` -- the component list -- has answers too variable to assert on: the same configuration returned 1 to 6 of 7 components. See research.md - [x] T005a [P] Assert at retrieval level that `complexes` was searched, in tests/retrievers/test_sync_async_equivalence.py -- parametrised over all five real collection names: each is reachable, and selecting it searches nothing else. Verified by sabotaging `resolve_collections` to ignore the selection, which fails all five. **Catches a plumbing failure, not a routing one**: a classifier that never chooses a collection is still undetectable, and needs a production probe rather than a test (research.md) +- [ ] T005b **The risk T005 was written for is still open.** T005a catches a plumbing failure; a classifier that never *chooses* a collection is undetectable by any deterministic test, and produces confident plausible answers rather than an error. Needs the routing distribution observed over real traffic, or a periodic probe asking known-collection questions and checking what was selected. **Do not treat Phase 2 as closing this** -- a register that reads "handled" is why nobody looks again - [x] T006 [P] Add an `ewas`-dependent question to `EXPECTATIONS` in src/evaluation/answer_sweep.py (TP53 UniProt P04637; verified by removal) - [x] T007 Assert at retrieval level that `reactions` was searched -- closed by the same parametrised test as T005a, in tests/retrievers/test_sync_async_equivalence.py. No answer-level question can guard it, because every reaction name also appears in `summations` - [x] T008 Verify each new question FAILS when its collection is removed from the bundle copy, and passes with it present; record the evidence in the PR (method established; two of four candidates survived it)