diff --git a/specs/010-search-page-answers/research.md b/specs/010-search-page-answers/research.md index 25b1802..d142fb0 100644 --- a/specs/010-search-page-answers/research.md +++ b/specs/010-search-page-answers/research.md @@ -169,3 +169,79 @@ expansion exists for the questions nobody wrote a test for. So this ships as `QUERY_EXPANSION_ALTERNATES` with the default at 4 -- a switch and a measurement, not a verdict. Trying 0 on beta, where the sweep and the routing probe both run on every deploy, is the cheap way to learn more. + +## T022 -- is 2s/10s reachable? Measured 2026-09-20 + +Re-measured because the recorded number has been wrong three times. Taken on +`astream_answer`, the path the endpoint uses. The two arms are **interleaved**, +alternating which goes first, and the number of queries actually asked is +recorded so "expansion off" is shown rather than assumed. n=25 per arm. + +| | default | expansion off | +|---|---|---| +| queries per retrieval | 5 | 1 | +| preprocess | 1.58s | 1.65s | +| retrieval | 3.15s | 0.29s | +| answer model (residual) | 0.84s | 0.79s | +| **first token, p50** | **5.57s** | **2.73s** | +| first token, min–max | 2.26s – 11.44s | 2.36s – 5.59s | +| **complete, p50** | **8.90s** | **7.16s** | + +### The plain answer + +**10s complete is met**, at 8.90s p50 today and 7.16s with expansion off. + +**2s to first token is not met**, either way. Disabling query expansion gets +p50 to 2.73s -- short of the target but the same order as it rather than +double. So FR-005a stays a target with a named blocker. + +### The tail is the finding, not the median + +The default path's **best** case is 2.26s and its worst is 11.44s -- a fivefold +spread. Disabling expansion barely moves the best case (2.36s) and halves the +worst (5.59s). + +So expansion does not make a typical request slow; it makes the slow requests +much slower. A target quoted at p50 would hide that entirely, and the panel +renders progressively, so the tail is the part a reader actually notices. Any +future latency requirement should be stated at a high percentile. + +Reported as min–max rather than p90: at n=25 a p90 is little more than the +second-highest sample, and the first attempt quoted one from n=15, where it +was barely more than the maximum. + +### A number I cannot reconcile + +The previous record has the answer model's time to first token at **6.1s**, +"the largest block left". This puts it at **0.84s**, with retrieval the +largest at 3.15s. + +My figure is a *residual* -- first token minus preprocess minus retrieval -- +so it absorbs anything unattributed, which biases it upward, not down. It is +also stable across both arms (0.84s and 0.79s) where the arms differ hugely in +retrieval, which is what a real per-call cost looks like rather than a +measurement artefact. + +They disagree by a factor of seven and only one can describe today's code. +Stated as unreconciled rather than quietly replacing the old number, because +"the answer model is the problem" has been steering this spec's priorities. + +### What the interleaving showed + +The first attempt ran the arms in sequence, default first, so warming would +have favoured expansion-off. Interleaving changed the conclusion not at all, +and the internal check for it is `preprocess`: it is 1.58s and 1.65s across +arms that differ tenfold in retrieval, and expansion cannot affect +preprocessing, so a sequencing confound would have shown up there and did not. + +### What would have to change, in measured order + +1. **Retrieval, 3.15s** -- the largest block, and 2.86s of it is query + expansion. `QUERY_EXPANSION_ALTERNATES=0` removes it today and is one + variable to revert. Its recall cost is unmeasured beyond thirteen tracked + questions (T020). +2. **Preprocessing, 1.58s** -- whether a search-page question needs all four + calls is still open (T020c). The two rounds removed the waiting, not the + calls. +3. **The answer model, 0.84s** -- not worth attention on these numbers, which + is exactly why the discrepancy above matters. diff --git a/specs/010-search-page-answers/tasks.md b/specs/010-search-page-answers/tasks.md index 463250c..bdda90a 100644 --- a/specs/010-search-page-answers/tasks.md +++ b/specs/010-search-page-answers/tasks.md @@ -61,7 +61,7 @@ state. - [x] T020 Measure query expansion, and make the count configurable. **The task's premise was wrong**: the cost is the expansion *call*, not the number of variants it returns. Measured over six tracked questions -- 4 alternates 1.27s expand + 1.22s retrieve; 2 alternates 1.26s + 0.65s; 0 alternates 0.00s + 0.31s. Trimming the count saves only fan-out; the cost disappears only at zero, where `answer-sweep` was 13/13 in 79s against ~150s. `QUERY_EXPANSION_ALTERNATES` sets it and **the default is unchanged**: thirteen questions show those answers do not need expansion, not that recall is unaffected in general - [ ] T020c Establish whether a search-page question needs all four preprocessing calls. The sequential half of this is answered and done (T019a): they run in two rounds and cost 2.6s at the median, not the ~16s recorded here, which never reproduced - [ ] T021 Land spec 009 collection routing and re-measure. **Still open — I marked this done on 2026-09-18 and was wrong.** What landed is *source* routing (`resolve_active_sources` picks reactome / userguide / live). Collection routing is selecting among the five collections *within* the reactome bundle, and it is not implemented: `QueryIntent` has no `collections` field, `resolve_collections` does not exist, and `retrieve_documents` still loops over every collection -- [ ] T022 Re-assess FR-005 against the result and say plainly whether 2s/10s is reachable +- [x] T022 Re-assessed 2026-09-20 on the served path, arms interleaved, n=25 each, with the query count asserted rather than assumed. **10s complete is met** (8.90s p50, 7.16s with expansion off); **2s to first token is not** (5.57s p50, 2.73s with expansion off), so FR-005a stays a target with a named blocker. The blocker changed: retrieval is the largest block at 3.15s and the answer model is 0.84s against a previously recorded 6.1s, recorded as unreconciled. The tail is the finding -- default first token ranges 2.26s to 11.44s, and expansion barely moves the best case while halving the worst, so a future target belongs at a high percentile - [x] T023 Propagate a real failure signal out of the live path. `answer_from_live_services` takes an optional `LiveReport` and records a tool exception **before** stringifying it into the model's context -- the last point at which it is still a fact rather than a paraphrase. The graph carries it out as `live_tool_failed`, and `answer_sweep` retries on that instead of matching prose. `TRANSIENT` and `_looks_transient` are gone, with a test asserting they do not come back: if prose matching returns it will return as a list of markers. The report is optional, so none of the twelve existing call sites changed. **The no-tools fallback deliberately does not set it**: `get_mcp_tools` remembers a failed start, so a retry could never succeed and marking it transient would buy an attempt guaranteed to fail. Pinned by a test, because the absence of the flag there reads like an oversight ## Phase 6: Handover