Skip to content

Explain the statistics with the reader's own numbers - #280

Merged
adamjohnwright merged 1 commit into
mainfrom
011-phase5-statistics
Sep 20, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
011-phase5-statistics

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Spec 011 Phase 5.

Fragility is computed, not requested

A hit resting on fewer than five matched entities is flagged per pathway, because a model handed a tiny p-value and a tiny count describes the p-value. That is the specific mistake the spec's own rationale names: "a pathway with 2 of 3 entities found is not strong evidence".

The threshold is on the count, not the ratio, and that is the point. 2 of 3 looks like a perfect hit by ratio and is nearly meaningless; 40 of 200 looks poor and is real evidence. The ratio is the misleading number, so the flag ignores it — pinned by a test using both.

The instruction also fixes what the significance language means: significant in the data is after correction, and a small p-value that does not survive correction is not a finding.

Two things from running it against a real result

It worked, and leaked a field name. The summary named the actual counts — "2 found out of 4 total", which is what US3 asks for — but wrote "these pathways are classified as fragile", handing the reader an internal label instead of the reason. The instruction now forbids repeating field names. Three further runs: no leaked label, real counts in each.

I nearly credited that fix wrongly. The assertion checking it failed, because it looked for the phrase in the source file, where the string is wrapped across two lines. So the improvement I saw in that run was model non-determinism, not my edit. The edit was fine; the check was wrong — and only running it three times distinguished the two.

Scope note

No per-pathway request parameter was added. US3's independent test implies one, the website has not asked for it, and every shown pathway already carries its own counts — so the explanation is per-pathway without new contract surface.

609 passed, 1 skipped; mypy over 151 files, ruff clean.

Process note on myself: I committed this once with 6 mypy errors, because | tail -1 swallowed the exit code in my check chain. Re-done with set -e and each check gated. That is a repeat of a mistake already on record.

🤖 Generated with Claude Code

Spec 011 Phase 5.

Fragility is computed in code rather than asked for in the prompt. A hit
resting on fewer than five matched entities is flagged per pathway, because a
model handed a tiny p-value and a tiny count describes the p-value -- which is
the specific mistake the spec's rationale names: "a pathway with 2 of 3
entities found is not strong evidence".

The threshold is on the count, not the ratio, and that is the whole point. 2
of 3 looks like a perfect hit by ratio and is nearly meaningless; 40 of 200
looks poor and is real evidence. The ratio is the misleading number here, so
the flag ignores it. Pinned by a test using both.

The instruction also fixes what the significance language means: `significant`
in the data is after correction, and a small p-value that does not survive
correction is not a finding.

Two things from running it against a real result rather than a stub.

It worked -- the summary names the actual counts, "2 found out of 4 total",
which is what US3 asks for -- but it wrote "these pathways are classified as
fragile", handing the reader one of our internal field names instead of the
reason. The instruction now forbids repeating field names and asks for the
explanation in the reader's terms. Three further runs, no leaked label, real
counts in each.

And I nearly credited that fix wrongly. The assertion checking it had failed,
because it looked for the phrase in the *source file* where the string is
wrapped across two lines -- so the improvement I saw in that run was model
non-determinism, not my edit. The edit was fine; the check was wrong, and only
running it three times distinguished the two.

No per-pathway request parameter was added. US3's independent test implies
one, the website has not asked for it, and every shown pathway already carries
its own counts -- so the explanation is per-pathway without new contract
surface.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 153c7be into main Sep 20, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the 011-phase5-statistics branch September 20, 2026 17:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant