diff --git a/specs/010-search-page-answers/research.md b/specs/010-search-page-answers/research.md index d142fb0..05c12c8 100644 --- a/specs/010-search-page-answers/research.md +++ b/specs/010-search-page-answers/research.md @@ -245,3 +245,87 @@ preprocessing, so a sequencing confound would have shown up there and did not. calls. 3. **The answer model, 0.84s** -- not worth attention on these numbers, which is exactly why the discrepancy above matters. + +## T020c -- does a search-page question need all four preprocessing calls? + +Measured 2026-09-20. The answer is yes for three of them, and "not for the +reason you would guess" for the fourth. + +### The rephraser is not a no-op without history + +It exists to resolve a follow-up against chat history, and the search page +sends one question on a fresh thread. The obvious conclusion is that it does +nothing there. Measured over the fifteen tracked questions with empty +history, it **changed 14 of 15**: + +- "Which release of Reactome is this?" -> "Which version of Reactome is + currently available?" +- "What does CDK5 phosphorylate in Alzheimer disease?" -> "What are the + substrates that CDK5 phosphorylates in the context of..." + +That is query normalisation rather than follow-up resolution: it feeds +retrieval and intent classification with different text than it was given. + +**"Changed" is not "improved", and the two facts here point opposite ways.** +It rewrites 14 of 15 questions, *and* the sweep passes without it -- so for +the tracked set those rewrites are not load-bearing. What the measurement +supports is that removing it is a retrieval-*input* change rather than a +plumbing one, and therefore needs measuring on something other than the +thirteen questions that already pass either way. It does not support the +claim that the rewrites are valuable. + +### With it bypassed, the sweep still passes -- and latency barely moves + +`answer-sweep` was 13/13 with the rephraser replaced by a pass-through +(precondition asserted: bypassed 13 times). + +Latency, both arms **interleaved in one window**, n=25 each: + +| | first token p50 | min | max | +|---|---|---|---| +| expansion off | 2.94s | 2.35s | 5.94s | +| expansion off + rephrase off | **2.83s** | 2.04s | 8.12s | + +About **0.11s**, which is the round arithmetic below rather than the ~0.8s the +call's own duration suggests. + +A first version of this compared 2.97s against a 2.73s taken in an earlier +run and reported "no better, slightly worse" -- the cross-window comparison +that the T022 review had just corrected, repeated here a few hours later. The +direction was wrong; the conclusion that the saving is ~0.1s rather than ~0.8s +was not, and is now better supported. + +### Why removing a 0.88s call saved 0.07s + +The arithmetic that predicted ~0.8s was wrong, and the error is worth +recording because it is easy to repeat. + +Preprocessing runs two rounds. Round one is `max(rephrase, language)` and +round two is `max(safety, intent)`. Median call times are rephrase 0.88s, +language 0.81s, safety 0.70s, intent 0.93s. So round one costs 0.88s with the +rephraser and **0.81s without it** -- language detection is nearly as slow, and +it was hiding behind the rephraser the whole time. + +Bypassing the call does not remove the round. The saving needs the rounds +**restructured**: with no rephrasing, `rephrased_input` is just `user_input`, +so safety and intent no longer depend on round one and all three calls can run +together -- 0.93s instead of 1.74s. That is the ~0.8s, and it requires a code +change this measurement did not make. + +### So, for FR-005a + +**2s to first token is not reachable by removing these two calls.** Expansion +off gets p50 to 2.94s and adding rephrase-off reaches 2.83s -- the best single +observation dipped to 2.04s, and the median did not. Restructuring preprocessing into one round is worth +about 0.8s more on paper, which would put p50 near 2s -- on paper, and against +a tail that neither change addresses. + +### What this does not establish + +Both levers pass the same thirteen questions, and that is now being asked to +carry a lot. Each alone keeps the sweep green; together they keep it green; +none of that shows recall is unaffected, and removing *two* recall mechanisms +on one thin evidence base compounds the risk rather than adding to the +confidence. The rephraser rewriting 14 of 15 questions is the concrete reason +to be careful: whatever those rewrites are worth, the sweep is not what +measures it. diff --git a/specs/010-search-page-answers/tasks.md b/specs/010-search-page-answers/tasks.md index bdda90a..6be6bf0 100644 --- a/specs/010-search-page-answers/tasks.md +++ b/specs/010-search-page-answers/tasks.md @@ -59,7 +59,7 @@ state. - [x] T019 Measure first-token and completion separately across the tracked questions; publish the distribution, not one question. Measured 2026-09-18, two runs each: first token p50 9.6s / p90 12.2s (n=26), completion p50 10.4s / p90 18.1s (n=30). The earlier "36.1s to first token" came from one question and does not reproduce (PR #238) - [x] T019a Run preprocessing in two rounds instead of four sequential calls in src/agent/profiles/react_to_me.py; the base class already overlapped, and this override discarded it (PR #238) - [x] T020 Measure query expansion, and make the count configurable. **The task's premise was wrong**: the cost is the expansion *call*, not the number of variants it returns. Measured over six tracked questions -- 4 alternates 1.27s expand + 1.22s retrieve; 2 alternates 1.26s + 0.65s; 0 alternates 0.00s + 0.31s. Trimming the count saves only fan-out; the cost disappears only at zero, where `answer-sweep` was 13/13 in 79s against ~150s. `QUERY_EXPANSION_ALTERNATES` sets it and **the default is unchanged**: thirteen questions show those answers do not need expansion, not that recall is unaffected in general -- [ ] T020c Establish whether a search-page question needs all four preprocessing calls. The sequential half of this is answered and done (T019a): they run in two rounds and cost 2.6s at the median, not the ~16s recorded here, which never reproduced +- [x] T020c Established 2026-09-20. **The rephraser is not a no-op without history** -- it changed 14 of 15 tracked questions, normalising queries rather than resolving follow-ups, so removing it is a retrieval-quality change. Bypassed, the sweep is 13/13 and first token p50 goes 2.94s to 2.83s (interleaved, n=25): about 0.11s. Removing a 0.88s call saved 0.07s because round one is `max(rephrase, language)` and language detection is 0.81s -- it was hiding behind the rephraser. The ~0.8s needs the **rounds restructured** so safety and intent stop waiting, which is a code change this did not make. 2s is not reachable by removing these two calls - [ ] T021 Land spec 009 collection routing and re-measure. **Still open — I marked this done on 2026-09-18 and was wrong.** What landed is *source* routing (`resolve_active_sources` picks reactome / userguide / live). Collection routing is selecting among the five collections *within* the reactome bundle, and it is not implemented: `QueryIntent` has no `collections` field, `resolve_collections` does not exist, and `retrieve_documents` still loops over every collection - [x] T022 Re-assessed 2026-09-20 on the served path, arms interleaved, n=25 each, with the query count asserted rather than assumed. **10s complete is met** (8.90s p50, 7.16s with expansion off); **2s to first token is not** (5.57s p50, 2.73s with expansion off), so FR-005a stays a target with a named blocker. The blocker changed: retrieval is the largest block at 3.15s and the answer model is 0.84s against a previously recorded 6.1s, recorded as unreconciled. The tail is the finding -- default first token ranges 2.26s to 11.44s, and expansion barely moves the best case while halving the worst, so a future target belongs at a high percentile - [x] T023 Propagate a real failure signal out of the live path. `answer_from_live_services` takes an optional `LiveReport` and records a tool exception **before** stringifying it into the model's context -- the last point at which it is still a fact rather than a paraphrase. The graph carries it out as `live_tool_failed`, and `answer_sweep` retries on that instead of matching prose. `TRANSIENT` and `_looks_transient` are gone, with a test asserting they do not come back: if prose matching returns it will return as a list of markers. The report is optional, so none of the twelve existing call sites changed. **The no-tools fallback deliberately does not set it**: `get_mcp_tools` remembers a failed start, so a retry could never succeed and marking it transient would buy an attempt guaranteed to fail. Pinned by a test, because the absence of the flag there reads like an oversight