Repository navigation
A fast local preflight by default; test totals counted at a release - #841
Merged
Merged
Conversation
npm run preflight runs the suite under coverage on a developer's machine beside whatever else it is doing, and there the 5 s default measured the machine, not the test: in four preflight runs on 2026-10-05, a test that takes 45 ms alone timed out at 5 s twice and a process-starting test once, each failing a 25-minute run that CI would have passed. CI's runners keep the default, so a genuinely slow test still fails there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…elease The preflight ran every check CI runs, one after another: 23 minutes on a desktop against CI's 11 in parallel, and twice whenever a branch added a test, because the committed truthbase carried the test totals and the only way to match them was to run the whole suite again and count. CI did the same: claims-alignment ran the suite a second time just to count it. - npm run preflight now runs, in a few minutes, the checks a branch fails CI on for no reason but a missed step: claims and renders, lint, the type checks, the dashboard's and the website's checks when the branch touches them, and the test files the branch adds or changes. CI runs the whole suite, in parallel. `--full` keeps the old run for a change wide enough to want it. - generate.mjs --check keeps the committed test totals unless --with-test-counts asks for them; claims-alignment no longer recounts. The release roll captures them, as it already did. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
Iris gate — 1 of 2 tripped
|
| Trace | Verdict | Basis | Rules, classes or missing inputs | Evidence |
|---|---|---|---|---|
b1b42bb540f9518c6209f8b835af5d14 |
failed | detector_veto + risk_over_loss |
no_pii, pii_leak, credential_leak | no_pii: AWS Access Key (output 45–65) |
| Verdict basis | Traces |
|---|---|
detector_veto |
2 |
clean |
1 |
Unjudged questions: task_completed (3), tool_use_correct (3) — a trace that did not carry what a rule needs.
tests/fixtures/ci-gate/traces.ndjson · 3 evaluated · dataset release-gate: 2 in the gate · exit 1 · what the bases mean
Iris gate — 1 stored, nothing tripped
|
| Verdict basis | Traces |
|---|---|
clean |
1 |
Unjudged questions: task_completed (1), tool_use_correct (1) — a trace that did not carry what a rule needs.
tests/fixtures/ci-gate/clean.ndjson · 1 evaluated · exit 0 · what the bases mean
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
Contributing to this repository had become slower than it needed to be. Three changes:
npm run preflight(and the pre-push hook)--fullkeeps the old run.claims.jsongenerate.mjs --check --with-test-counts; counted at a release, as the release roll already does.claims-alignmentno longer runs the suite a second timeWhy
Measured on 2026-10-05:
website-build-scope.test.tsalone runs a few hundred git commands. So the fast preflight does not run the whole suite, and not every test a change reaches either: a change to a module every test imports reaches all of it. CI runs the suite.Tests
tests/preflight-mirrors-ci.test.ts: every pull-request job is still a--fullstep or CI-only. The fast mode runs the branch's own test files, never the whole suite or the builds; the dashboard's and the website's checks only when the branch touches them; and--fullruns none of the fast-only steps.npm run preflight(fast) passed on this commit in 74 s.🤖 Generated with Claude Code