Removing a 0.88s call saved 0.07s, and the reason is the useful part - #275
Merged
Merged
Conversation
T020c. Three of the four preprocessing calls are plainly needed. The fourth, the rephraser, exists to resolve a follow-up against chat history -- and the search page sends one question on a fresh thread, so the obvious conclusion is that it does nothing there. It changed 14 of the 15 tracked questions with empty history. Not follow-up resolution: query normalisation, feeding retrieval and intent classification. "Which release of Reactome is this?" becomes "Which version of Reactome is currently available?". Removing it is a retrieval-quality change, not a plumbing one. Bypassed, the sweep is 13/13. But first token p50 is 2.97s with this and query expansion both off, against 2.73s with expansion off alone -- no better, and inside the noise. The arithmetic that predicted ~0.8s was wrong, which is the part worth keeping. Preprocessing runs two rounds: `max(rephrase, language)` then `max(safety, intent)`. Rephrase is 0.88s and language detection 0.81s, so round one costs 0.81s without the rephraser instead of 0.88s. Language was hiding behind it the whole time. Bypassing a call does not remove a round. The saving needs the rounds restructured -- with no rephrasing, safety and intent no longer depend on round one and all three can run together, 0.93s instead of 1.74s. That is a code change this measurement did not make, and it is where the 0.8s actually is. So 2s to first token is not reachable by removing these two calls, and the thirteen tracked questions are now carrying more weight than they can. Each lever alone keeps the sweep green and so does the pair, but none of that shows recall is unaffected -- and removing two recall mechanisms on one thin evidence base compounds the risk rather than adding confidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… fixed Two corrections to T020c. The latency claim compared 2.97s against a 2.73s taken in an earlier run, and reported "no better, slightly worse". That is the cross-window comparison the T022 review corrected a few hours earlier, repeated here. Re-run with the arms interleaved in one window, n=25 each: expansion-off 2.94s, both-off 2.83s. So bypassing the rephraser saves about 0.11s -- the direction of my claim was wrong, and the conclusion that the saving is ~0.1s rather than ~0.8s is now better supported than it was by the comparison that appeared to prove it. And "it changed 14 of 15 questions, so it is doing real work" overstates what was measured. Changed is not improved. The sweep passes without it, so for the tracked set those rewrites are not load-bearing, and the two facts point in opposite directions rather than both supporting caution. What the measurement supports is narrower: removing it changes retrieval *input*, so it needs measuring on something other than the thirteen questions that pass either way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
T020c — the last open item in spec 010. Documentation only.
The rephraser is not a no-op without history
It exists to resolve a follow-up against chat history, and the search page sends one question on a fresh thread. The obvious conclusion is that it does nothing there. Measured over the fifteen tracked questions with empty history, it changed 14 of 15:
"Which release of Reactome is this?"→"Which version of Reactome is currently available?""What does CDK5 phosphorylate in Alzheimer disease?"→"What are the substrates that CDK5 phosphorylates in the context of..."That is query normalisation, not follow-up resolution — real work feeding retrieval and intent classification. Removing it is a retrieval-quality change, not a plumbing one.
Bypassed, the sweep passes and latency barely moves
answer-sweep13/13 with a pass-through rephraser (precondition asserted: bypassed 13 times). But first token p50 2.97s with this and query expansion both off, against 2.73s with expansion off alone — no better, inside the noise, and the tail returns (max 11.31s).Why removing a 0.88s call saved 0.07s
The arithmetic that predicted ~0.8s was wrong, and the error is easy to repeat.
Preprocessing runs two rounds:
max(rephrase, language)thenmax(safety, intent). Median call times are rephrase 0.88s, language 0.81s, safety 0.70s, intent 0.93s. So round one costs 0.81s without the rephraser instead of 0.88s — language detection is nearly as slow and was hiding behind it.Bypassing a call does not remove a round. The saving needs the rounds restructured: with no rephrasing,
rephrased_inputis justuser_input, so safety and intent stop depending on round one and all three run together — 0.93s instead of 1.74s. That is a code change this measurement did not make, and it is where the 0.8s actually is.For FR-005a
2s to first token is not reachable by removing these two calls. Restructuring preprocessing into one round is worth ~0.8s more on paper, which would put p50 near 2s — on paper, and against a tail neither change addresses.
What this does not establish
Both levers pass the same thirteen questions, and that set is now carrying a lot. Each alone keeps the sweep green; together they keep it green; none of it shows recall is unaffected. Removing two recall mechanisms on one thin evidence base compounds the risk rather than adding confidence — and the rephraser rewriting 14 of 15 questions is the concrete reason to be careful about what the sweep does not measure.
🤖 Generated with Claude Code