Put the measurement checks in the code rather than in my memory - #276
Merged
Merged
Conversation
Three mistakes have recurred in this repository's measurements, each more than once, and knowing the lesson has not prevented repeating it within hours -- the cross-window comparison was corrected in one PR and repeated in the next. `src/evaluation/ab.py` makes each one structurally hard. Arms alternate. There is no mode that runs them in sequence, because the second arm otherwise inherits a warm process and whatever the API is doing that minute. That has twice produced a result in the direction the author was hoping for, and once showed a real improvement as a regression. Every arm declares a `precondition`, checked after each sample, and the summary refuses to draw a conclusion when one fails. A prompt that asked for "exactly 1" alternate returned four; a ContextVar set around a call was overwritten by the node inside it. Both produced ordinary-looking numbers that were measuring the baseline twice. And the summary reports p50 with min and max, never a percentile the sample cannot support. A p90 was quoted from fifteen samples, where it is about the second-highest value. The configuration is applied before every sample rather than once per arm, because a setting applied once and mutated in between fails the same way as never applying it. Each test recreates a mistake actually made here, and both guards were verified by sabotage: removing the alternation fails the alternation test, ignoring preconditions fails the refusal test. Exercised on the real graph too, not only on stubs -- expansion on 6.22s against off 2.86s, precondition held. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
adamjohnwright
force-pushed
the
evaluation-ab-harness
branch
from
September 20, 2026 04:42
fd9f458 to
0f5e7db
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three mistakes have recurred in this repository's measurements, each more than once. Knowing the lesson has not prevented repeating it — the cross-window comparison was corrected in one PR and repeated in the next, hours later.
src/evaluation/ab.pymakes each one structurally hard rather than a thing to remember.What it enforces
Arms alternate. There is no mode that runs them in sequence. The second arm otherwise inherits a warm process, a warm connection, and whatever the API is doing that minute — that has twice produced a result in the direction the author was hoping for, and once showed a real improvement as a 2.5s regression.
Every arm declares a
precondition, checked after each sample, and the summary refuses to draw a conclusion when one fails. A prompt asking for "exactly 1" alternate returned four; a ContextVar set around a call was overwritten by the node inside it. Both produced ordinary-looking numbers that were quietly measuring the baseline twice.p50 with min and max, never a percentile the sample cannot support. A p90 was quoted from fifteen samples, where it is about the second-highest value.
Configuration is applied before every sample, not once per arm — a setting applied once and mutated in between fails the same way as never applying it, and looks identical.
How it is verified
Each test recreates a mistake actually made here, not a hypothetical one. Both guards were checked by sabotage:
And it is exercised on the real graph, not only on stubs: expansion on
p50 6.22sagainst offp50 2.86s, precondition held,trustworthy: True.What it does not do
It does not make a measurement correct. It removes three specific ways of being confidently wrong, all three of which cost real rework today. Attribution errors, residuals absorbing unmeasured time, and thin question sets are all still available.
591 passed, 1 skipped; mypy over 151 files, ruff clean.
🤖 Generated with Claude Code