Skip to content

CI: a shorter critical path, a timeout on every job, one run per PR - #845

Merged
irparent merged 1 commit into
mainfrom
ci/faster-critical-path
Oct 5, 2026
Merged

irparent merged 1 commit into
mainfrom
ci/faster-critical-path

Conversation

@irparent

@irparent irparent commented Oct 5, 2026

Copy link
Copy Markdown
Member

What changes

Measured on main's recent CI runs (every job's p95 is under 6 minutes, and the run took 11.6 to 15.2 minutes):

Before Now
Job ordering 12 jobs waited on lint-and-typecheck, test or build without using anything they produced. e2e waited for build, then rebuilt everything itself, so it started 6–9 minutes into the run and finished last every time only the two pack jobs keep needs:, because they pass artifacts; everything else starts at once
Job timeouts none, so a hang held a runner and kept a PR pending for up to 6 hours 20 minutes on every CI job
Superseded runs a new push left the previous run going ci, claims-alignment, codeql, lighthouse and publish-python keep one run per PR; main and tags are never cancelled
Docker layer cache every PR wrote its own, which only that PR can read; the cache stood at 10.7 GB against GitHub's 10 GB limit only main writes it; PRs read it
Lighthouse (not required) every PR and push, with a seed step ending || true only when the dashboard, its server, the seed data or the budget changes; a failed seed fails the job

Nothing that gates a merge changes: the same jobs run, with the same names.

Tests

  • actionlint (the pinned image CI uses) passes on every workflow.
  • The workflow tests pass: preflight-mirrors-ci, workflows-build-order, release-workflow-consistency, ci-gate-action, site-says-what-the-code-does, required-checks-documented.
  • npm run preflight passed on this commit.

🤖 Generated with Claude Code

Measured on main's recent runs (job p95 at most 6 minutes):

- Twelve CI jobs waited on lint-and-typecheck, test or build without
  using anything they produced; e2e waited for build and then rebuilt
  everything itself, so it started 6 to 9 minutes into the run and
  finished last every time. Only the two pack jobs pass artifacts; they
  keep their needs, every other job starts at once.
- No job had a timeout, so a hang (this repo's recurring bug class is
  processes that never exit) held a runner and kept a PR pending for up
  to 6 hours. Every CI job now stops at 20 minutes.
- No PR workflow cancelled the run a new push replaced. ci,
  claims-alignment, codeql, lighthouse and publish-python now keep one
  run per pull request; main and tags are never cancelled.
- Every PR wrote its own Docker layer cache, which only that PR can
  read; the repository's cache stood at 10.7 GB against a 10 GB limit.
  Only main writes it now.
- Lighthouse (not a required check) runs when the dashboard, its server,
  the seed data or its budget change, and a failed seed fails it instead
  of scoring an empty dashboard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 5, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
website Ignored Ignored Oct 5, 2026 9:55pm UTC

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Iris gate — 1 of 2 tripped --fail-on detector_veto

iris-eval ingest: 3 stored, 1 tripped --fail-on detector_veto (2 of 3 evaluated in dataset "release-gate")

Trace Verdict Basis Rules, classes or missing inputs Evidence
d2dd6d74f5ccc21c05dfe2d4abeae042 failed detector_veto + risk_over_loss no_pii, pii_leak, credential_leak no_pii: AWS Access Key (output 45–65)
Verdict basis Traces
detector_veto 2
clean 1

Unjudged questions: task_completed (3), tool_use_correct (3) — a trace that did not carry what a rule needs.

tests/fixtures/ci-gate/traces.ndjson · 3 evaluated · dataset release-gate: 2 in the gate · exit 1 · what the bases mean

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Iris gate — 1 stored, nothing tripped --fail-on any

iris-eval ingest: 1 stored, 0 tripped --fail-on any

Verdict basis Traces
clean 1

Unjudged questions: task_completed (1), tool_use_correct (1) — a trace that did not carry what a rule needs.

tests/fixtures/ci-gate/clean.ndjson · 1 evaluated · exit 0 · what the bases mean

@irparent
irparent merged commit 8c0e0d3 into main Oct 5, 2026
74 checks passed
@irparent
irparent deleted the ci/faster-critical-path branch October 5, 2026 22:16
irparent added a commit that referenced this pull request Oct 5, 2026
…its own

#845 gave ci, claims-alignment, codeql, lighthouse and publish-python one
concurrency group per pull request, cancelling a run its PR replaced. On
main the group was the branch, so a merge queued behind a running one
replaced the merge queued before it: two of main's commits on 10-05 never
got their CI run. A pull request keeps its group; every other run (a push
to main, a tag, a schedule) is grouped by its own run id, so none is
replaced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
irparent added a commit that referenced this pull request Oct 5, 2026
The 20-minute limit every job gained in #845 cancelled e2e on a pull
request whose tests were passing: the browser download took 976 s against
about a minute normally (Playwright's download server, not the change).
The browsers are now cached per Playwright version, so most runs download
nothing, and e2e, the one job that depends on someone else's server for a
large download, gets 40 minutes; every other job keeps 20.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
irparent added a commit that referenced this pull request Oct 5, 2026
…its own

#845 gave ci, claims-alignment, codeql, lighthouse and publish-python one
concurrency group per pull request, cancelling a run its PR replaced. On
main the group was the branch, so a merge queued behind a running one
replaced the merge queued before it: two of main's commits on 10-05 never
got their CI run. A pull request keeps its group; every other run (a push
to main, a tag, a schedule) is grouped by its own run id, so none is
replaced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
irparent added a commit that referenced this pull request Oct 5, 2026
…ns never replace each other (#852)

* ci: e2e keeps the Playwright browsers in a cache, and gets 40 minutes

The 20-minute limit every job gained in #845 cancelled e2e on a pull
request whose tests were passing: the browser download took 976 s against
about a minute normally (Playwright's download server, not the change).
The browsers are now cached per Playwright version, so most runs download
nothing, and e2e, the one job that depends on someone else's server for a
large download, gets 40 minutes; every other job keeps 20.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: every run that is not a pull request gets a concurrency group of its own

#845 gave ci, claims-alignment, codeql, lighthouse and publish-python one
concurrency group per pull request, cancelling a run its PR replaced. On
main the group was the branch, so a merge queued behind a running one
replaced the merge queued before it: two of main's commits on 10-05 never
got their CI run. A pull request keeps its group; every other run (a push
to main, a tag, a schedule) is grouped by its own run id, so none is
replaced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant