diff --git a/specs/010-search-page-answers/contracts/answer_endpoint.md b/specs/010-search-page-answers/contracts/answer_endpoint.md index e4e58ee..a893a40 100644 --- a/specs/010-search-page-answers/contracts/answer_endpoint.md +++ b/specs/010-search-page-answers/contracts/answer_endpoint.md @@ -221,9 +221,20 @@ A naive measurement reports 3.0s to first token. That token is the rephraser's. | retrieval finishing to the answer's first token | 6.1 | The largest block is no longer retrieval or preprocessing: it is the answer model's -own time to first token with a retrieved context. Collection routing -([009](../../009-collection-routing/spec.md)) already did most of the work on -retrieval -- the 12.5s figure recorded here before it landed no longer reproduces. +own time to first token with a retrieved context. + +**A correction to an earlier version of this file**, which said collection routing +had landed and explained the improvement. It has not. What landed is *source* +routing -- a question goes to the Reactome bundle, the user guide, or a live +lookup -- while [009](../../009-collection-routing/spec.md), selecting among the +collections *within* the Reactome bundle, is still unimplemented; retrieval +searches all five. + +So the honest position on why the earlier 12.5s and ~36s figures no longer +reproduce is that we do not fully know. Source routing accounts for the fast +user-guide answers (3.3s to first token). Running preprocessing in two rounds +accounts for about 2.5s. The rest may simply be that those figures came from one +question and were never representative. One measurement worth keeping in view for anyone optimising this: the async retrieval path is **not faster than the sync one** here -- 12.5s against 10.9s on the diff --git a/specs/010-search-page-answers/tasks.md b/specs/010-search-page-answers/tasks.md index 0f4b154..28d2e5c 100644 --- a/specs/010-search-page-answers/tasks.md +++ b/specs/010-search-page-answers/tasks.md @@ -60,7 +60,7 @@ state. - [x] T019a Run preprocessing in two rounds instead of four sequential calls in src/agent/profiles/react_to_me.py; the base class already overlapped, and this override discarded it (PR #238) - [ ] T020 Reduce query expansion from 5 variants, measuring recall with bin/retrieval_baseline — its own call plus a 5x retrieval fan-out - [ ] T020c Establish whether a search-page question needs all four preprocessing calls. The sequential half of this is answered and done (T019a): they run in two rounds and cost 2.6s at the median, not the ~16s recorded here, which never reproduced -- [x] T021 Land spec 009 collection routing and re-measure — landed and live (`state["active_sources"][0]` selects at retrieval time), and re-measured 2026-09-18: it is the main reason first-token fell from the ~36s once recorded to 9.6s +- [ ] T021 Land spec 009 collection routing and re-measure. **Still open — I marked this done on 2026-09-18 and was wrong.** What landed is *source* routing (`resolve_active_sources` picks reactome / userguide / live). Collection routing is selecting among the five collections *within* the reactome bundle, and it is not implemented: `QueryIntent` has no `collections` field, `resolve_collections` does not exist, and `retrieve_documents` still loops over every collection - [ ] T022 Re-assess FR-005 against the result and say plainly whether 2s/10s is reachable ## Phase 6: Handover