Measure query expansion; the cost is the call, not the variants - #272
Merged
Merged
Conversation
T020 said "reduce query expansion from 5 variants". The premise was wrong. Measured over six tracked questions: four alternates cost 1.27s to generate and 1.22s to retrieve for; two cost 1.26s and 0.65s; zero cost nothing and 0.31s. The expansion call costs about 1.27s whatever it returns, so trimming the count saves only fan-out -- roughly 0.57s of a 2.49s stage. The cost disappears only by not making the call. With expansion off, `answer-sweep` passed 13/13 in 79s against about 150s. That is the largest single latency lever found so far and it bears on T022. The count is now `QUERY_EXPANSION_ALTERNATES`, enforced in code rather than asked for in the prompt, and **the default is unchanged at 4**. Thirteen questions establish that those answers do not need expansion; they do not establish that recall is unaffected in general, and expansion exists for the questions nobody wrote a test for. A switch and a measurement, not a verdict. At zero the call is skipped rather than made and discarded -- it is the larger half of the cost, and making it anyway would keep the expense while losing the benefit, while looking identical in every other measurement. A bad value falls back to the default loudly, because a typo silently disabling a recall mechanism would look like nothing at all downstream. How it was measured matters here. The first attempt varied the count by editing the prompt to ask for "exactly N". The model obeyed at 2 and ignored 1 and 0, producing four either way, so those rows re-measured the baseline -- caught only because the number of queries actually asked was recorded next to the timings. That is also why the implementation truncates rather than asks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
adamjohnwright
force-pushed
the
010-t020-query-expansion
branch
from
September 20, 2026 02:29
43d87df to
622889c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
T020 said "reduce query expansion from 5 variants". The premise was wrong, and the measurement is the useful part.
What it costs
Six tracked questions, median:
The expansion call costs ~1.27s whatever it returns. Trimming four variants to two saves fan-out only — about 0.57s of a 2.49s stage. The cost disappears only by not making the call, which is a different change from the one the task described.
With expansion off,
answer-sweeppassed 13/13 in 79s against roughly 150s. That is the largest single latency lever found so far, and it bears directly on T022.What ships
QUERY_EXPANSION_ALTERNATES, enforced in code rather than asked for in the prompt, with the default unchanged at 4.Thirteen tracked questions establish that those answers do not need expansion. They do not establish that recall is unaffected in general, and expansion exists for the questions nobody wrote a test for. So this is a switch and a measurement, not a verdict — the call is yours.
Two details that are behaviour, not tidiness:
How it was measured
The first attempt varied the count by editing the prompt to ask for "exactly N". The model obeyed at 2 and ignored 1 and 0, producing four either way — so those rows silently re-measured the baseline. Visible only because the number of queries actually asked was recorded next to the timings. That is also why the implementation truncates rather than asks.
Document overlap is not quality: 71% at zero says three-quarters of the retrieved set is unchanged, not that the changed quarter did not matter.
Suggested next step
Try
QUERY_EXPANSION_ALTERNATES=0on beta, where the sweep and the routing probe both run on every deploy. That is the cheap way to learn whether the recall loss is real, and it is one environment variable to revert.582 passed, 1 skipped; mypy over 149 files, ruff clean.
🤖 Generated with Claude Code