Repository navigation
Say plainly whether 2s/10s is reachable: one yes, one no - #274
Merged
Merged
Conversation
T022, measured on `astream_answer` rather than `ainvoke`, five questions by three runs per arm, with wall clock to the first token recorded independently of the phase timings so the breakdown has to account for the total. 10s complete is met: 9.64s at p50, 7.04s with query expansion disabled. 2s to first token is not, either way. 5.27s on the default path and 2.81s without expansion -- short of the target but the same order as it rather than double. So FR-005a stays a target with a named blocker. What changed is which blocker. The spec records the answer model's time to first token at 6.1s and calls it the largest block left. This measurement puts it at 0.66s, with retrieval the largest at 2.84s. I cannot reconcile the two. My figure is a residual -- first token minus preprocess minus retrieval -- so it absorbs unattributed time and if anything reads high. They disagree by a factor of nine and only one can describe today's code. Recorded as unreconciled rather than quietly replacing the old number, because "the answer model is the problem" has been steering this spec's priorities. And the variance is worse than the median suggests: p90 first token is 10.57s against a 5.27s p50, collapsing to 3.81s without expansion. A target quoted at p50 hides that one request in ten waits twice the median, on a panel that renders progressively. Any future requirement should be stated at p90. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A measurement PR is worth only what the measurement is worth, and mine had three faults. The arms ran in sequence, default first, so any warming favoured expansion-off -- the arm whose case I was making. They interleave now, alternating per run. The number of queries actually asked was never recorded, so "expansion off" was assumed rather than shown. That is the precondition check I have insisted on twice today and skipped here. Asserted now: 5 queries against 1. And p90 came from fifteen samples, where it is barely more than the maximum. n=25 per arm now, and the spread is reported as min-max rather than a percentile the sample cannot support. The conclusions hold. 10s complete met at 8.90s; 2s to first token not met at 5.57s, or 2.73s with expansion off. Two things sharpened. The tail is the finding rather than the median: the default path ranges 2.26s to 11.44s, and disabling expansion barely moves the best case while halving the worst. Expansion does not make a typical request slow, it makes the slow ones much slower -- which a p50 target would hide entirely. And the answer-model residual is now 0.84s and 0.79s across arms that differ tenfold in retrieval. A stable per-call cost rather than an artefact, which strengthens rather than excuses the unreconciled 6.1s on record. The interleaving has its own check: `preprocess` is 1.58s and 1.65s across those same arms, and expansion cannot affect preprocessing, so a sequencing confound would have surfaced there and did not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
T022. Measured on
astream_answer— the path the endpoint uses — over five questions × three runs per arm, with wall clock to the first token recorded independently of the phase timings so the breakdown has to account for the total rather than being assumed to.The plain answer
10s complete: met. 9.64s today, 7.04s with expansion off.
2s to first token: not met, either way, and not close on the default path. Disabling query expansion reaches 2.81s — short of the target but the same order as it rather than double. So FR-005a stays a target with a named blocker.
A number I cannot reconcile
The spec records the answer model's time to first token at 6.1s, "the largest block left". This measurement puts it at 0.66s, with retrieval the largest at 2.84s.
I can offer the mechanism but not the proof. My figure is a residual — first token minus preprocess minus retrieval — so it absorbs anything unattributed, which biases it upward, not down. The two disagree by a factor of nine and only one can describe today's code.
Recorded as unreconciled rather than quietly replacing the old number, because "the answer model is the problem" has been steering this spec's priorities and now looks wrong. Someone should settle it before acting on either.
The variance matters more than the median
p90 first token is 10.57s against a 5.27s median — one request in ten waits twice the median, on a panel that renders progressively, so the wait is visible. Expansion off collapses p90 to 3.81s. Any future latency requirement should be stated at p90, because a p50 target hides exactly the requests that feel broken.
What would have to change, in measured order
QUERY_EXPANSION_ALTERNATES=0removes it today; recall cost unmeasured beyond thirteen questions (T020).Documentation only; no code changes.
🤖 Generated with Claude Code