Skip to content

Say plainly whether 2s/10s is reachable: one yes, one no - #274

Merged
adamjohnwright merged 2 commits into
mainfrom
010-t022-latency-assessment
Sep 20, 2026
Merged

adamjohnwright merged 2 commits into
mainfrom
010-t022-latency-assessment

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

T022. Measured on astream_answer — the path the endpoint uses — over five questions × three runs per arm, with wall clock to the first token recorded independently of the phase timings so the breakdown has to account for the total rather than being assumed to.

default expansion off
preprocess 1.77s 1.65s
retrieval 2.84s 0.26s
answer model (residual) 0.66s 0.90s
first token, p50 5.27s 2.81s
first token, p90 10.57s 3.81s
complete, p50 9.64s 7.04s

The plain answer

10s complete: met. 9.64s today, 7.04s with expansion off.

2s to first token: not met, either way, and not close on the default path. Disabling query expansion reaches 2.81s — short of the target but the same order as it rather than double. So FR-005a stays a target with a named blocker.

A number I cannot reconcile

The spec records the answer model's time to first token at 6.1s, "the largest block left". This measurement puts it at 0.66s, with retrieval the largest at 2.84s.

I can offer the mechanism but not the proof. My figure is a residual — first token minus preprocess minus retrieval — so it absorbs anything unattributed, which biases it upward, not down. The two disagree by a factor of nine and only one can describe today's code.

Recorded as unreconciled rather than quietly replacing the old number, because "the answer model is the problem" has been steering this spec's priorities and now looks wrong. Someone should settle it before acting on either.

The variance matters more than the median

p90 first token is 10.57s against a 5.27s median — one request in ten waits twice the median, on a panel that renders progressively, so the wait is visible. Expansion off collapses p90 to 3.81s. Any future latency requirement should be stated at p90, because a p50 target hides exactly the requests that feel broken.

What would have to change, in measured order

  1. Retrieval, 2.84s — largest block, and 2.58s of it is query expansion. QUERY_EXPANSION_ALTERNATES=0 removes it today; recall cost unmeasured beyond thirteen questions (T020).
  2. Preprocessing, 1.77s — whether a search-page question needs all four calls is still open (T020c).
  3. The answer model, 0.66s — no longer worth attention on these numbers, which is why the discrepancy above matters.

Documentation only; no code changes.

🤖 Generated with Claude Code

adamjohnwright and others added 2 commits September 20, 2026 03:09
T022, measured on `astream_answer` rather than `ainvoke`, five questions by
three runs per arm, with wall clock to the first token recorded independently
of the phase timings so the breakdown has to account for the total.

10s complete is met: 9.64s at p50, 7.04s with query expansion disabled.

2s to first token is not, either way. 5.27s on the default path and 2.81s
without expansion -- short of the target but the same order as it rather than
double. So FR-005a stays a target with a named blocker.

What changed is which blocker. The spec records the answer model's time to
first token at 6.1s and calls it the largest block left. This measurement puts
it at 0.66s, with retrieval the largest at 2.84s. I cannot reconcile the two.
My figure is a residual -- first token minus preprocess minus retrieval -- so
it absorbs unattributed time and if anything reads high. They disagree by a
factor of nine and only one can describe today's code. Recorded as
unreconciled rather than quietly replacing the old number, because "the answer
model is the problem" has been steering this spec's priorities.

And the variance is worse than the median suggests: p90 first token is 10.57s
against a 5.27s p50, collapsing to 3.81s without expansion. A target quoted at
p50 hides that one request in ten waits twice the median, on a panel that
renders progressively. Any future requirement should be stated at p90.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A measurement PR is worth only what the measurement is worth, and mine had
three faults.

The arms ran in sequence, default first, so any warming favoured
expansion-off -- the arm whose case I was making. They interleave now,
alternating per run.

The number of queries actually asked was never recorded, so "expansion off"
was assumed rather than shown. That is the precondition check I have insisted
on twice today and skipped here. Asserted now: 5 queries against 1.

And p90 came from fifteen samples, where it is barely more than the maximum.
n=25 per arm now, and the spread is reported as min-max rather than a
percentile the sample cannot support.

The conclusions hold. 10s complete met at 8.90s; 2s to first token not met at
5.57s, or 2.73s with expansion off. Two things sharpened.

The tail is the finding rather than the median: the default path ranges 2.26s
to 11.44s, and disabling expansion barely moves the best case while halving
the worst. Expansion does not make a typical request slow, it makes the slow
ones much slower -- which a p50 target would hide entirely.

And the answer-model residual is now 0.84s and 0.79s across arms that differ
tenfold in retrieval. A stable per-call cost rather than an artefact, which
strengthens rather than excuses the unreconciled 6.1s on record.

The interleaving has its own check: `preprocess` is 1.58s and 1.65s across
those same arms, and expansion cannot affect preprocessing, so a sequencing
confound would have surfaced there and did not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit ea01db9 into main Sep 20, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the 010-t022-latency-assessment branch September 20, 2026 03:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant