Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 14 additions & 3 deletions specs/010-search-page-answers/contracts/answer_endpoint.md
Original file line number Diff line number Diff line change
Expand Up @@ -221,9 +221,20 @@ A naive measurement reports 3.0s to first token. That token is the rephraser's.
| retrieval finishing to the answer's first token | 6.1 |

The largest block is no longer retrieval or preprocessing: it is the answer model's
own time to first token with a retrieved context. Collection routing
([009](../../009-collection-routing/spec.md)) already did most of the work on
retrieval -- the 12.5s figure recorded here before it landed no longer reproduces.
own time to first token with a retrieved context.

**A correction to an earlier version of this file**, which said collection routing
had landed and explained the improvement. It has not. What landed is *source*
routing -- a question goes to the Reactome bundle, the user guide, or a live
lookup -- while [009](../../009-collection-routing/spec.md), selecting among the
collections *within* the Reactome bundle, is still unimplemented; retrieval
searches all five.

So the honest position on why the earlier 12.5s and ~36s figures no longer
reproduce is that we do not fully know. Source routing accounts for the fast
user-guide answers (3.3s to first token). Running preprocessing in two rounds
accounts for about 2.5s. The rest may simply be that those figures came from one
question and were never representative.

One measurement worth keeping in view for anyone optimising this: the async
retrieval path is **not faster than the sync one** here -- 12.5s against 10.9s on the
Expand Down
2 changes: 1 addition & 1 deletion specs/010-search-page-answers/tasks.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,7 @@ state.
- [x] T019a Run preprocessing in two rounds instead of four sequential calls in src/agent/profiles/react_to_me.py; the base class already overlapped, and this override discarded it (PR #238)
- [ ] T020 Reduce query expansion from 5 variants, measuring recall with bin/retrieval_baseline — its own call plus a 5x retrieval fan-out
- [ ] T020c Establish whether a search-page question needs all four preprocessing calls. The sequential half of this is answered and done (T019a): they run in two rounds and cost 2.6s at the median, not the ~16s recorded here, which never reproduced
- [x] T021 Land spec 009 collection routing and re-measure — landed and live (`state["active_sources"][0]` selects at retrieval time), and re-measured 2026-09-18: it is the main reason first-token fell from the ~36s once recorded to 9.6s
- [ ] T021 Land spec 009 collection routing and re-measure. **Still open — I marked this done on 2026-09-18 and was wrong.** What landed is *source* routing (`resolve_active_sources` picks reactome / userguide / live). Collection routing is selecting among the five collections *within* the reactome bundle, and it is not implemented: `QueryIntent` has no `collections` field, `resolve_collections` does not exist, and `retrieve_documents` still loops over every collection
- [ ] T022 Re-assess FR-005 against the result and say plainly whether 2s/10s is reachable

## Phase 6: Handover
Expand Down
Loading