From fb77087c3633d7e3837c4d33ecbec767b8258b65 Mon Sep 17 00:00:00 2001 From: Adam Wright Date: Sun, 20 Sep 2026 03:47:25 +0000 Subject: [PATCH 1/2] Removing a 0.88s call saved 0.07s, and the reason is the useful part T020c. Three of the four preprocessing calls are plainly needed. The fourth, the rephraser, exists to resolve a follow-up against chat history -- and the search page sends one question on a fresh thread, so the obvious conclusion is that it does nothing there. It changed 14 of the 15 tracked questions with empty history. Not follow-up resolution: query normalisation, feeding retrieval and intent classification. "Which release of Reactome is this?" becomes "Which version of Reactome is currently available?". Removing it is a retrieval-quality change, not a plumbing one. Bypassed, the sweep is 13/13. But first token p50 is 2.97s with this and query expansion both off, against 2.73s with expansion off alone -- no better, and inside the noise. The arithmetic that predicted ~0.8s was wrong, which is the part worth keeping. Preprocessing runs two rounds: `max(rephrase, language)` then `max(safety, intent)`. Rephrase is 0.88s and language detection 0.81s, so round one costs 0.81s without the rephraser instead of 0.88s. Language was hiding behind it the whole time. Bypassing a call does not remove a round. The saving needs the rounds restructured -- with no rephrasing, safety and intent no longer depend on round one and all three can run together, 0.93s instead of 1.74s. That is a code change this measurement did not make, and it is where the 0.8s actually is. So 2s to first token is not reachable by removing these two calls, and the thirteen tracked questions are now carrying more weight than they can. Each lever alone keeps the sweep green and so does the pair, but none of that shows recall is unaffected -- and removing two recall mechanisms on one thin evidence base compounds the risk rather than adding confidence. Co-Authored-By: Claude Opus 5 --- specs/010-search-page-answers/research.md | 63 +++++++++++++++++++++++ specs/010-search-page-answers/tasks.md | 2 +- 2 files changed, 64 insertions(+), 1 deletion(-) diff --git a/specs/010-search-page-answers/research.md b/specs/010-search-page-answers/research.md index d142fb0..b9716d5 100644 --- a/specs/010-search-page-answers/research.md +++ b/specs/010-search-page-answers/research.md @@ -245,3 +245,66 @@ preprocessing, so a sequencing confound would have shown up there and did not. calls. 3. **The answer model, 0.84s** -- not worth attention on these numbers, which is exactly why the discrepancy above matters. + +## T020c -- does a search-page question need all four preprocessing calls? + +Measured 2026-09-20. The answer is yes for three of them, and "not for the +reason you would guess" for the fourth. + +### The rephraser is not a no-op without history + +It exists to resolve a follow-up against chat history, and the search page +sends one question on a fresh thread. The obvious conclusion is that it does +nothing there. Measured over the fifteen tracked questions with empty +history, it **changed 14 of 15**: + +- "Which release of Reactome is this?" -> "Which version of Reactome is + currently available?" +- "What does CDK5 phosphorylate in Alzheimer disease?" -> "What are the + substrates that CDK5 phosphorylates in the context of..." + +That is query normalisation, not follow-up resolution -- real work that feeds +retrieval and intent classification. Removing it is a retrieval-quality +change, not a plumbing one. + +### With it bypassed, the sweep still passes -- and latency barely moves + +`answer-sweep` was 13/13 with the rephraser replaced by a pass-through +(precondition asserted: bypassed 13 times). But first token p50 went to +**2.97s with both this and query expansion disabled, against 2.73s with +expansion disabled alone.** No better, and inside the noise. + +### Why removing a 0.88s call saved 0.07s + +The arithmetic that predicted ~0.8s was wrong, and the error is worth +recording because it is easy to repeat. + +Preprocessing runs two rounds. Round one is `max(rephrase, language)` and +round two is `max(safety, intent)`. Median call times are rephrase 0.88s, +language 0.81s, safety 0.70s, intent 0.93s. So round one costs 0.88s with the +rephraser and **0.81s without it** -- language detection is nearly as slow, and +it was hiding behind the rephraser the whole time. + +Bypassing the call does not remove the round. The saving needs the rounds +**restructured**: with no rephrasing, `rephrased_input` is just `user_input`, +so safety and intent no longer depend on round one and all three calls can run +together -- 0.93s instead of 1.74s. That is the ~0.8s, and it requires a code +change this measurement did not make. + +### So, for FR-005a + +**2s to first token is not reachable by removing these two calls.** Expansion +off gets p50 to 2.73s; adding rephrase-off does not improve it, and the tail +returns (max 11.31s). Restructuring preprocessing into one round is worth +about 0.8s more on paper, which would put p50 near 2s -- on paper, and against +a tail that neither change addresses. + +### What this does not establish + +Both levers pass the same thirteen questions, and that is now being asked to +carry a lot. Each alone keeps the sweep green; together they keep it green; +none of that shows recall is unaffected, and removing *two* recall mechanisms +on one thin evidence base compounds the risk rather than adding to the +confidence. The rephraser rewriting 14 of 15 questions is the concrete reason +to be careful: whatever those rewrites are worth, the sweep is not what +measures it. diff --git a/specs/010-search-page-answers/tasks.md b/specs/010-search-page-answers/tasks.md index bdda90a..50361c1 100644 --- a/specs/010-search-page-answers/tasks.md +++ b/specs/010-search-page-answers/tasks.md @@ -59,7 +59,7 @@ state. - [x] T019 Measure first-token and completion separately across the tracked questions; publish the distribution, not one question. Measured 2026-09-18, two runs each: first token p50 9.6s / p90 12.2s (n=26), completion p50 10.4s / p90 18.1s (n=30). The earlier "36.1s to first token" came from one question and does not reproduce (PR #238) - [x] T019a Run preprocessing in two rounds instead of four sequential calls in src/agent/profiles/react_to_me.py; the base class already overlapped, and this override discarded it (PR #238) - [x] T020 Measure query expansion, and make the count configurable. **The task's premise was wrong**: the cost is the expansion *call*, not the number of variants it returns. Measured over six tracked questions -- 4 alternates 1.27s expand + 1.22s retrieve; 2 alternates 1.26s + 0.65s; 0 alternates 0.00s + 0.31s. Trimming the count saves only fan-out; the cost disappears only at zero, where `answer-sweep` was 13/13 in 79s against ~150s. `QUERY_EXPANSION_ALTERNATES` sets it and **the default is unchanged**: thirteen questions show those answers do not need expansion, not that recall is unaffected in general -- [ ] T020c Establish whether a search-page question needs all four preprocessing calls. The sequential half of this is answered and done (T019a): they run in two rounds and cost 2.6s at the median, not the ~16s recorded here, which never reproduced +- [x] T020c Established 2026-09-20. **The rephraser is not a no-op without history** -- it changed 14 of 15 tracked questions, normalising queries rather than resolving follow-ups, so removing it is a retrieval-quality change. Bypassed, the sweep is 13/13 but first token p50 is 2.97s against 2.73s with expansion off alone: no gain. Removing a 0.88s call saved 0.07s because round one is `max(rephrase, language)` and language detection is 0.81s -- it was hiding behind the rephraser. The ~0.8s needs the **rounds restructured** so safety and intent stop waiting, which is a code change this did not make. 2s is not reachable by removing these two calls - [ ] T021 Land spec 009 collection routing and re-measure. **Still open — I marked this done on 2026-09-18 and was wrong.** What landed is *source* routing (`resolve_active_sources` picks reactome / userguide / live). Collection routing is selecting among the five collections *within* the reactome bundle, and it is not implemented: `QueryIntent` has no `collections` field, `resolve_collections` does not exist, and `retrieve_documents` still loops over every collection - [x] T022 Re-assessed 2026-09-20 on the served path, arms interleaved, n=25 each, with the query count asserted rather than assumed. **10s complete is met** (8.90s p50, 7.16s with expansion off); **2s to first token is not** (5.57s p50, 2.73s with expansion off), so FR-005a stays a target with a named blocker. The blocker changed: retrieval is the largest block at 3.15s and the answer model is 0.84s against a previously recorded 6.1s, recorded as unreconciled. The tail is the finding -- default first token ranges 2.26s to 11.44s, and expansion barely moves the best case while halving the worst, so a future target belongs at a high percentile - [x] T023 Propagate a real failure signal out of the live path. `answer_from_live_services` takes an optional `LiveReport` and records a tool exception **before** stringifying it into the model's context -- the last point at which it is still a fact rather than a paraphrase. The graph carries it out as `live_tool_failed`, and `answer_sweep` retries on that instead of matching prose. `TRANSIENT` and `_looks_transient` are gone, with a test asserting they do not come back: if prose matching returns it will return as a list of markers. The report is optional, so none of the twelve existing call sites changed. **The no-tools fallback deliberately does not set it**: `get_mcp_tools` remembers a failed start, so a retry could never succeed and marking it transient would buy an attempt guaranteed to fail. Pinned by a test, because the absence of the flag there reads like an oversight From 3a4bc94d9eeff91a59382e30c0bec9cef55d7daf Mon Sep 17 00:00:00 2001 From: Adam Wright Date: Sun, 20 Sep 2026 04:31:02 +0000 Subject: [PATCH 2/2] Adversarial review: I repeated the cross-window comparison I had just fixed Two corrections to T020c. The latency claim compared 2.97s against a 2.73s taken in an earlier run, and reported "no better, slightly worse". That is the cross-window comparison the T022 review corrected a few hours earlier, repeated here. Re-run with the arms interleaved in one window, n=25 each: expansion-off 2.94s, both-off 2.83s. So bypassing the rephraser saves about 0.11s -- the direction of my claim was wrong, and the conclusion that the saving is ~0.1s rather than ~0.8s is now better supported than it was by the comparison that appeared to prove it. And "it changed 14 of 15 questions, so it is doing real work" overstates what was measured. Changed is not improved. The sweep passes without it, so for the tracked set those rewrites are not load-bearing, and the two facts point in opposite directions rather than both supporting caution. What the measurement supports is narrower: removing it changes retrieval *input*, so it needs measuring on something other than the thirteen questions that pass either way. Co-Authored-By: Claude Opus 5 --- specs/010-search-page-answers/research.md | 37 ++++++++++++++++++----- specs/010-search-page-answers/tasks.md | 2 +- 2 files changed, 30 insertions(+), 9 deletions(-) diff --git a/specs/010-search-page-answers/research.md b/specs/010-search-page-answers/research.md index b9716d5..05c12c8 100644 --- a/specs/010-search-page-answers/research.md +++ b/specs/010-search-page-answers/research.md @@ -263,16 +263,37 @@ history, it **changed 14 of 15**: - "What does CDK5 phosphorylate in Alzheimer disease?" -> "What are the substrates that CDK5 phosphorylates in the context of..." -That is query normalisation, not follow-up resolution -- real work that feeds -retrieval and intent classification. Removing it is a retrieval-quality -change, not a plumbing one. +That is query normalisation rather than follow-up resolution: it feeds +retrieval and intent classification with different text than it was given. + +**"Changed" is not "improved", and the two facts here point opposite ways.** +It rewrites 14 of 15 questions, *and* the sweep passes without it -- so for +the tracked set those rewrites are not load-bearing. What the measurement +supports is that removing it is a retrieval-*input* change rather than a +plumbing one, and therefore needs measuring on something other than the +thirteen questions that already pass either way. It does not support the +claim that the rewrites are valuable. ### With it bypassed, the sweep still passes -- and latency barely moves `answer-sweep` was 13/13 with the rephraser replaced by a pass-through -(precondition asserted: bypassed 13 times). But first token p50 went to -**2.97s with both this and query expansion disabled, against 2.73s with -expansion disabled alone.** No better, and inside the noise. +(precondition asserted: bypassed 13 times). + +Latency, both arms **interleaved in one window**, n=25 each: + +| | first token p50 | min | max | +|---|---|---|---| +| expansion off | 2.94s | 2.35s | 5.94s | +| expansion off + rephrase off | **2.83s** | 2.04s | 8.12s | + +About **0.11s**, which is the round arithmetic below rather than the ~0.8s the +call's own duration suggests. + +A first version of this compared 2.97s against a 2.73s taken in an earlier +run and reported "no better, slightly worse" -- the cross-window comparison +that the T022 review had just corrected, repeated here a few hours later. The +direction was wrong; the conclusion that the saving is ~0.1s rather than ~0.8s +was not, and is now better supported. ### Why removing a 0.88s call saved 0.07s @@ -294,8 +315,8 @@ change this measurement did not make. ### So, for FR-005a **2s to first token is not reachable by removing these two calls.** Expansion -off gets p50 to 2.73s; adding rephrase-off does not improve it, and the tail -returns (max 11.31s). Restructuring preprocessing into one round is worth +off gets p50 to 2.94s and adding rephrase-off reaches 2.83s -- the best single +observation dipped to 2.04s, and the median did not. Restructuring preprocessing into one round is worth about 0.8s more on paper, which would put p50 near 2s -- on paper, and against a tail that neither change addresses. diff --git a/specs/010-search-page-answers/tasks.md b/specs/010-search-page-answers/tasks.md index 50361c1..6be6bf0 100644 --- a/specs/010-search-page-answers/tasks.md +++ b/specs/010-search-page-answers/tasks.md @@ -59,7 +59,7 @@ state. - [x] T019 Measure first-token and completion separately across the tracked questions; publish the distribution, not one question. Measured 2026-09-18, two runs each: first token p50 9.6s / p90 12.2s (n=26), completion p50 10.4s / p90 18.1s (n=30). The earlier "36.1s to first token" came from one question and does not reproduce (PR #238) - [x] T019a Run preprocessing in two rounds instead of four sequential calls in src/agent/profiles/react_to_me.py; the base class already overlapped, and this override discarded it (PR #238) - [x] T020 Measure query expansion, and make the count configurable. **The task's premise was wrong**: the cost is the expansion *call*, not the number of variants it returns. Measured over six tracked questions -- 4 alternates 1.27s expand + 1.22s retrieve; 2 alternates 1.26s + 0.65s; 0 alternates 0.00s + 0.31s. Trimming the count saves only fan-out; the cost disappears only at zero, where `answer-sweep` was 13/13 in 79s against ~150s. `QUERY_EXPANSION_ALTERNATES` sets it and **the default is unchanged**: thirteen questions show those answers do not need expansion, not that recall is unaffected in general -- [x] T020c Established 2026-09-20. **The rephraser is not a no-op without history** -- it changed 14 of 15 tracked questions, normalising queries rather than resolving follow-ups, so removing it is a retrieval-quality change. Bypassed, the sweep is 13/13 but first token p50 is 2.97s against 2.73s with expansion off alone: no gain. Removing a 0.88s call saved 0.07s because round one is `max(rephrase, language)` and language detection is 0.81s -- it was hiding behind the rephraser. The ~0.8s needs the **rounds restructured** so safety and intent stop waiting, which is a code change this did not make. 2s is not reachable by removing these two calls +- [x] T020c Established 2026-09-20. **The rephraser is not a no-op without history** -- it changed 14 of 15 tracked questions, normalising queries rather than resolving follow-ups, so removing it is a retrieval-quality change. Bypassed, the sweep is 13/13 and first token p50 goes 2.94s to 2.83s (interleaved, n=25): about 0.11s. Removing a 0.88s call saved 0.07s because round one is `max(rephrase, language)` and language detection is 0.81s -- it was hiding behind the rephraser. The ~0.8s needs the **rounds restructured** so safety and intent stop waiting, which is a code change this did not make. 2s is not reachable by removing these two calls - [ ] T021 Land spec 009 collection routing and re-measure. **Still open — I marked this done on 2026-09-18 and was wrong.** What landed is *source* routing (`resolve_active_sources` picks reactome / userguide / live). Collection routing is selecting among the five collections *within* the reactome bundle, and it is not implemented: `QueryIntent` has no `collections` field, `resolve_collections` does not exist, and `retrieve_documents` still loops over every collection - [x] T022 Re-assessed 2026-09-20 on the served path, arms interleaved, n=25 each, with the query count asserted rather than assumed. **10s complete is met** (8.90s p50, 7.16s with expansion off); **2s to first token is not** (5.57s p50, 2.73s with expansion off), so FR-005a stays a target with a named blocker. The blocker changed: retrieval is the largest block at 3.15s and the answer model is 0.84s against a previously recorded 6.1s, recorded as unreconciled. The tail is the finding -- default first token ranges 2.26s to 11.44s, and expansion barely moves the best case while halving the worst, so a future target belongs at a high percentile - [x] T023 Propagate a real failure signal out of the live path. `answer_from_live_services` takes an optional `LiveReport` and records a tool exception **before** stringifying it into the model's context -- the last point at which it is still a fact rather than a paraphrase. The graph carries it out as `live_tool_failed`, and `answer_sweep` retries on that instead of matching prose. `TRANSIENT` and `_looks_transient` are gone, with a test asserting they do not come back: if prose matching returns it will return as a list of markers. The report is optional, so none of the twelve existing call sites changed. **The no-tools fallback deliberately does not set it**: `get_mcp_tools` remembers a failed start, so a retry could never succeed and marking it transient would buy an attempt guaranteed to fail. Pinned by a test, because the absence of the flag there reads like an oversight