Skip to content

Measure query expansion; the cost is the call, not the variants - #272

Merged
adamjohnwright merged 1 commit into
mainfrom
010-t020-query-expansion
Sep 20, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
010-t020-query-expansion

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

T020 said "reduce query expansion from 5 variants". The premise was wrong, and the measurement is the useful part.

What it costs

Six tracked questions, median:

alternates expansion call retrieval total documents kept
4 (default) 1.27s 1.22s 2.49s baseline
2 1.26s 0.65s 1.92s 84%
1 1.38s 0.57s 1.95s 77%
0 0.00s 0.31s 0.31s 71%

The expansion call costs ~1.27s whatever it returns. Trimming four variants to two saves fan-out only — about 0.57s of a 2.49s stage. The cost disappears only by not making the call, which is a different change from the one the task described.

With expansion off, answer-sweep passed 13/13 in 79s against roughly 150s. That is the largest single latency lever found so far, and it bears directly on T022.

What ships

QUERY_EXPANSION_ALTERNATES, enforced in code rather than asked for in the prompt, with the default unchanged at 4.

Thirteen tracked questions establish that those answers do not need expansion. They do not establish that recall is unaffected in general, and expansion exists for the questions nobody wrote a test for. So this is a switch and a measurement, not a verdict — the call is yours.

Two details that are behaviour, not tidiness:

  • At zero the call is skipped, not made and discarded. It is the larger half of the cost; making it anyway would keep the expense, lose the benefit, and look identical in every other measurement.
  • A bad value falls back loudly. A typo silently disabling a recall mechanism would look like nothing at all downstream.

How it was measured

The first attempt varied the count by editing the prompt to ask for "exactly N". The model obeyed at 2 and ignored 1 and 0, producing four either way — so those rows silently re-measured the baseline. Visible only because the number of queries actually asked was recorded next to the timings. That is also why the implementation truncates rather than asks.

Document overlap is not quality: 71% at zero says three-quarters of the retrieved set is unchanged, not that the changed quarter did not matter.

Suggested next step

Try QUERY_EXPANSION_ALTERNATES=0 on beta, where the sweep and the routing probe both run on every deploy. That is the cheap way to learn whether the recall loss is real, and it is one environment variable to revert.

582 passed, 1 skipped; mypy over 149 files, ruff clean.

🤖 Generated with Claude Code

T020 said "reduce query expansion from 5 variants". The premise was wrong.

Measured over six tracked questions: four alternates cost 1.27s to generate
and 1.22s to retrieve for; two cost 1.26s and 0.65s; zero cost nothing and
0.31s. The expansion call costs about 1.27s whatever it returns, so trimming
the count saves only fan-out -- roughly 0.57s of a 2.49s stage. The cost
disappears only by not making the call.

With expansion off, `answer-sweep` passed 13/13 in 79s against about 150s.
That is the largest single latency lever found so far and it bears on T022.

The count is now `QUERY_EXPANSION_ALTERNATES`, enforced in code rather than
asked for in the prompt, and **the default is unchanged at 4**. Thirteen
questions establish that those answers do not need expansion; they do not
establish that recall is unaffected in general, and expansion exists for the
questions nobody wrote a test for. A switch and a measurement, not a verdict.

At zero the call is skipped rather than made and discarded -- it is the larger
half of the cost, and making it anyway would keep the expense while losing the
benefit, while looking identical in every other measurement. A bad value falls
back to the default loudly, because a typo silently disabling a recall
mechanism would look like nothing at all downstream.

How it was measured matters here. The first attempt varied the count by
editing the prompt to ask for "exactly N". The model obeyed at 2 and ignored
1 and 0, producing four either way, so those rows re-measured the baseline --
caught only because the number of queries actually asked was recorded next to
the timings. That is also why the implementation truncates rather than asks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright force-pushed the 010-t020-query-expansion branch from 43d87df to 622889c Compare September 20, 2026 02:29
@adamjohnwright
adamjohnwright merged commit 3042371 into main Sep 20, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the 010-t020-query-expansion branch September 20, 2026 02:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant