Explain the statistics with the reader's own numbers - #280
Merged
Merged
Conversation
Spec 011 Phase 5. Fragility is computed in code rather than asked for in the prompt. A hit resting on fewer than five matched entities is flagged per pathway, because a model handed a tiny p-value and a tiny count describes the p-value -- which is the specific mistake the spec's rationale names: "a pathway with 2 of 3 entities found is not strong evidence". The threshold is on the count, not the ratio, and that is the whole point. 2 of 3 looks like a perfect hit by ratio and is nearly meaningless; 40 of 200 looks poor and is real evidence. The ratio is the misleading number here, so the flag ignores it. Pinned by a test using both. The instruction also fixes what the significance language means: `significant` in the data is after correction, and a small p-value that does not survive correction is not a finding. Two things from running it against a real result rather than a stub. It worked -- the summary names the actual counts, "2 found out of 4 total", which is what US3 asks for -- but it wrote "these pathways are classified as fragile", handing the reader one of our internal field names instead of the reason. The instruction now forbids repeating field names and asks for the explanation in the reader's terms. Three further runs, no leaked label, real counts in each. And I nearly credited that fix wrongly. The assertion checking it had failed, because it looked for the phrase in the *source file* where the string is wrapped across two lines -- so the improvement I saw in that run was model non-determinism, not my edit. The edit was fine; the check was wrong, and only running it three times distinguished the two. No per-pathway request parameter was added. US3's independent test implies one, the website has not asked for it, and every shown pathway already carries its own counts -- so the explanation is per-pathway without new contract surface. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Spec 011 Phase 5.
Fragility is computed, not requested
A hit resting on fewer than five matched entities is flagged per pathway, because a model handed a tiny p-value and a tiny count describes the p-value. That is the specific mistake the spec's own rationale names: "a pathway with 2 of 3 entities found is not strong evidence".
The threshold is on the count, not the ratio, and that is the point. 2 of 3 looks like a perfect hit by ratio and is nearly meaningless; 40 of 200 looks poor and is real evidence. The ratio is the misleading number, so the flag ignores it — pinned by a test using both.
The instruction also fixes what the significance language means:
significantin the data is after correction, and a small p-value that does not survive correction is not a finding.Two things from running it against a real result
It worked, and leaked a field name. The summary named the actual counts — "2 found out of 4 total", which is what US3 asks for — but wrote "these pathways are classified as fragile", handing the reader an internal label instead of the reason. The instruction now forbids repeating field names. Three further runs: no leaked label, real counts in each.
I nearly credited that fix wrongly. The assertion checking it failed, because it looked for the phrase in the source file, where the string is wrapped across two lines. So the improvement I saw in that run was model non-determinism, not my edit. The edit was fine; the check was wrong — and only running it three times distinguished the two.
Scope note
No per-pathway request parameter was added. US3's independent test implies one, the website has not asked for it, and every shown pathway already carries its own counts — so the explanation is per-pathway without new contract surface.
609 passed, 1 skipped; mypy over 151 files, ruff clean.
Process note on myself: I committed this once with 6 mypy errors, because
| tail -1swallowed the exit code in my check chain. Re-done withset -eand each check gated. That is a repeat of a mistake already on record.🤖 Generated with Claude Code