Skip to content

Put the measurement checks in the code rather than in my memory - #276

Merged
adamjohnwright merged 1 commit into
mainfrom
evaluation-ab-harness
Sep 20, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
evaluation-ab-harness

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Three mistakes have recurred in this repository's measurements, each more than once. Knowing the lesson has not prevented repeating it — the cross-window comparison was corrected in one PR and repeated in the next, hours later.

src/evaluation/ab.py makes each one structurally hard rather than a thing to remember.

What it enforces

Arms alternate. There is no mode that runs them in sequence. The second arm otherwise inherits a warm process, a warm connection, and whatever the API is doing that minute — that has twice produced a result in the direction the author was hoping for, and once showed a real improvement as a 2.5s regression.

Every arm declares a precondition, checked after each sample, and the summary refuses to draw a conclusion when one fails. A prompt asking for "exactly 1" alternate returned four; a ContextVar set around a call was overwritten by the node inside it. Both produced ordinary-looking numbers that were quietly measuring the baseline twice.

p50 with min and max, never a percentile the sample cannot support. A p90 was quoted from fifteen samples, where it is about the second-highest value.

Configuration is applied before every sample, not once per arm — a setting applied once and mutated in between fails the same way as never applying it, and looks identical.

How it is verified

Each test recreates a mistake actually made here, not a hypothetical one. Both guards were checked by sabotage:

  • removing the alternation → the alternation test fails
  • ignoring preconditions → the refusal test fails

And it is exercised on the real graph, not only on stubs: expansion on p50 6.22s against off p50 2.86s, precondition held, trustworthy: True.

What it does not do

It does not make a measurement correct. It removes three specific ways of being confidently wrong, all three of which cost real rework today. Attribution errors, residuals absorbing unmeasured time, and thin question sets are all still available.

591 passed, 1 skipped; mypy over 151 files, ruff clean.

🤖 Generated with Claude Code

Three mistakes have recurred in this repository's measurements, each more
than once, and knowing the lesson has not prevented repeating it within hours
-- the cross-window comparison was corrected in one PR and repeated in the
next.

`src/evaluation/ab.py` makes each one structurally hard.

Arms alternate. There is no mode that runs them in sequence, because the
second arm otherwise inherits a warm process and whatever the API is doing
that minute. That has twice produced a result in the direction the author was
hoping for, and once showed a real improvement as a regression.

Every arm declares a `precondition`, checked after each sample, and the
summary refuses to draw a conclusion when one fails. A prompt that asked for
"exactly 1" alternate returned four; a ContextVar set around a call was
overwritten by the node inside it. Both produced ordinary-looking numbers
that were measuring the baseline twice.

And the summary reports p50 with min and max, never a percentile the sample
cannot support. A p90 was quoted from fifteen samples, where it is about the
second-highest value.

The configuration is applied before every sample rather than once per arm,
because a setting applied once and mutated in between fails the same way as
never applying it.

Each test recreates a mistake actually made here, and both guards were
verified by sabotage: removing the alternation fails the alternation test,
ignoring preconditions fails the refusal test. Exercised on the real graph
too, not only on stubs -- expansion on 6.22s against off 2.86s, precondition
held.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit b6b8f07 into main Sep 20, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the evaluation-ab-harness branch September 20, 2026 11:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant