Skip to content

runner: eliminate idle CPU and bound reconciliation/JIT churn #408

Description

@donbeave

Problem

On 2026-08-26, Sentry consumed roughly 600–700% CPU while docker ps showed no containers and no user jobs running. The host had about 9–10 active velnor-runner daemon processes, each daemon main thread using approximately 50–84% CPU. Memory was normal (~9.5 GiB / 62.6 GiB).

This is control-plane work, not job execution.

Observed Sentry symptoms:

  • repeated registration lost messages;
  • failed JIT registration cleanup/delete;
  • repeated JIT reconfiguration attempts;
  • broker session HTTP 409 conflicts;
  • broker poll HTTP 404 failures;
  • stale GitHub registrations and timed-out unassigned jobs;
  • many idle slot and waiter/job processes despite no Docker containers;
  • doctor reporting degraded/failed fleet state while daemon health could still report ready.

Root cause

The current topology multiplies idle control work:

ready slot
  -> separate slot process
  -> separate waiter/job process
  -> separate broker session and poll loop
  -> registration/session drift
  -> JIT cleanup/re-registration churn
  -> repeated controller reconciliation and durable journal writes
  -> high CPU with zero Docker jobs

The strongest deterministic hot path is repeated reconciliation and journal write amplification. Every controller reconciles every two seconds, repeatedly applies proof/routing/capacity observations, and performs durable SQLite work. The per-slot waiter/job processes multiply this work and each owns independent broker/session recovery.

Relevant code:

  • crates/velnor-runner/src/node/controller.rs:232-275 — unconditional two-second reconciliation.
  • crates/velnor-runner/src/node/controller.rs:339-425 — per-cycle proof, registration, executor, session, routing, and journal work.
  • crates/velnor-runner/src/node/controller.rs:803-842 — waiter for every ready slot.
  • crates/velnor-runner/src/node/controller.rs:1002-1045 — waiter as separate velnor-runner job process.
  • crates/velnor-runner/src/node/job.rs:48-70 — waiter runs the full broker runner lifecycle.
  • crates/velnor-runner/src/runner.rs:1208-1215 — one OS process per slot, including surge capacity.
  • crates/velnor-runner/src/runner.rs:1411-1495 — lost registration triggers cleanup, fresh JIT configuration, and another slot cycle.
  • crates/velnor-runner/src/runner.rs:2666-2746 — independent idle broker polling loop.
  • crates/velnor-runner/src/runner.rs:3189-3281 — broker retry state is local to each runner.
  • crates/velnor-runner/src/protocol.rs:1841-1867 — broker 404 enters the error path.
  • crates/velnor-control/src/journal.rs:932-999 — each accepted event performs durable SQLite work.

Architecture decision

Use Tokio to consolidate idle control-plane concurrency, but do not immediately collapse the whole server into one address space.

The current normative Node Architecture v2 requires guardian, per-scope controller, per-ready-slot process, and transient job worker (docs/vision.md:82-91, docs/roadmap.md:31-35, docs/mission.md:50-55). This is a deliberate failure-containment boundary.

Target architecture:

guardian
  └── controller per scope
        ├── bounded async broker/session manager
        ├── event-driven reconciliation
        ├── isolated slot lifecycle
        └── transient job process → Docker or microVM execution

A single node-wide process is a later experiment only if it proves equivalent containment for panic, deadlock, protocol stall, CPU spin, OOM, and restart. One scope’s failure must never take down unrelated scopes or active jobs.

Implementation plan

Phase 0 — establish attribution and budgets

  • Add reconcile cycle count, duration p50/p95/p99, and overlap detection.
  • Add events/sec and accepted/rejected/no-op counts by event kind.
  • Add SQLite transaction count/time, lock wait, WAL size, and journal growth.
  • Add broker request count/status/latency and JIT create/delete count/status/latency.
  • Add registration/session generation, retry streak, next retry, and quarantine state.
  • Add CPU attribution by phase: journal, filesystem, GitHub, broker, child supervision.
  • Add process counts by role: daemon, controller, slot, waiter, job.
  • Add jobs, idle_slots, request rates, event rates, and resource status to doctor/health.
  • Reproduce zero-job/high-CPU behavior in a deterministic local or integration harness.
  • Record a fixed-hardware baseline before behavior changes.

Phase 1 — remove journal write amplification

  • Make PermitReserved, ExecutorProven, SessionLive, routing, and dependency observations no-op when unchanged.
  • Emit durable events only for meaningful state transitions.
  • Collect changed events per reconcile cycle and commit with one apply_many() transaction.
  • Preserve synchronous=FULL and crash durability.
  • Separate high-frequency heartbeat liveness from durable event history.
  • Keep atomic heartbeat files and validate them in the controller.
  • Journal only heartbeat transitions or bounded periodic summaries.
  • Prove replayed state equals materialized state after injected crash points.

Phase 2 — eliminate idle waiter processes

  • Remove the invariant that every ready idle slot starts a velnor-runner job waiter.
  • Keep idle broker polling in a bounded controller-owned Tokio session manager.
  • Start a transient job worker only after broker assignment.
  • Ensure zero jobs creates zero idle job-worker processes.
  • Ensure one assignment starts exactly one job worker.
  • Preserve completion, cancellation, cleanup, and generation fencing.
  • Preserve Docker and microVM isolation for execution.

Phase 3 — centralize broker and registration recovery

  • Implement one recovery authority per scope; sibling slots must not recreate the same identity independently.
  • Model lifecycle explicitly: Healthy → Missing/Conflict → Quarantined → Recreate → Registered.
  • Distinguish empty 202, 401, 403, 404, 409, timeout/reset, malformed response, rate limit, and registration disappearance.
  • Fence cleanup and mutation by registration/session generation.
  • Deduplicate concurrent recovery attempts.
  • Apply capped exponential backoff with jitter and a scope retry budget.
  • Quarantine repeatedly failing slots instead of immediate JIT churn.
  • Surface degraded/not-ready health while recovery is quarantined.
  • Ensure one scope’s API failure cannot amplify another scope’s retries.

Phase 4 — bound reconciliation

  • Replace unconditional full work every two seconds with event-triggered reconciliation plus a slow safety watchdog.
  • Bound work per cycle and prevent overlapping reconciliation.
  • Cache immutable/configuration observations between cycles.
  • Keep watchdog proof separate from useful-capacity/resource safety.
  • Make READY=1 and health distinguish control-cycle liveness from schedulable capacity.
  • Alert on sustained jobs=0 + high CPU.
  • Alert on repeated identical events and registration/JIT churn.

Phase 5 — tests and failure isolation

  • Unit-test no-op event suppression and batched journal commits.
  • Fault-inject broker empty response, 401, 403, 404, 409, timeout, malformed payload, rate limit, delete failure, and JIT failure.
  • Assert one registration loss causes one coordinated recovery path.
  • Assert stale generations cannot mutate newer sessions or registrations.
  • Assert retry deadlines and quarantine are honored.
  • Kill one slot; prove siblings remain alive.
  • Restart one controller; prove unrelated scopes remain alive.
  • Kill/block one broker task; prove other slots remain schedulable.
  • Fail one scope’s API path; prove other scopes continue normally.
  • Restart a controller during an active job; prove the job is preserved or enters explicit recovery.
  • Drain; prove idle polling stops while active jobs finish within the systemd bound.
  • Run all Rust tests with cargo nextest run.

Phase 6 — fixture and performance proof

  • Leave velnor-actions-fixture unchanged.
  • Run deterministic unit and fault-injection tests.
  • Run a multi-scope, zero-job idle soak for at least 15 minutes.
  • Verify idle CPU and broker/JIT request budgets.
  • Verify durable no-op events are zero and WAL growth is bounded.
  • Verify idle resource cost does not scale linearly with slot count.
  • Run fixture readiness and smoke tests.
  • Run GitHub-hosted/Velnor lane parity checks.
  • Compare conclusions, steps, logs, outputs, artifacts, caches, timing, and resource evidence.

Initial gates; calibrate only from the recorded baseline:

Signal Required gate
Zero-job CPU agreed fixed budget, initially ≤5% per scope
Idle CPU scaling ≤2× from 1 to 16 slots
Reconcile overlap 0
Stable-state JIT create/delete 0
Durable no-op events 0
Idle job workers 0
Retry behavior bounded by backoff and budget
Active-job p95 regression ≤5% versus baseline
WAL growth bounded by retention policy

Phase 7 — canary and release

  • Build and publish the candidate through signed Debian APT.
  • Drain and upgrade one isolated Sentry scope only.
  • Keep unrelated scopes on the previous version.
  • Observe idle soak and all resource/churn budgets.
  • Run one fixture job and one forced broker-fault sequence.
  • Prove active-job preservation, sibling-scope isolation, and no duplicate registrations.
  • Promote only after every gate passes.
  • Document exact rollback package/version and drain procedure.
  • Prove rollback through signed APT preserves active jobs and prevents stale-generation mutation.
  • Repeat idle soak and fixture proof after rollback validation.

Ready definition

Close this issue only when all are true:

  • CPU attribution is measured and the regression is reproduced.
  • Idle journal/reconciliation amplification is removed.
  • Idle waiter processes are gone; workers spawn only after assignment.
  • Broker/session recovery is coordinated, generation-fenced, bounded, and observable.
  • Health/doctor expose useful capacity, resource safety, and churn—not only control-loop liveness.
  • Existing process-isolation guarantees remain intact or have stronger equivalent proof.
  • Zero-job idle budgets pass for 15+ minutes.
  • Broker/JIT fault-injection tests pass without retry storms.
  • Fixture parity and smoke tests pass without fixture changes.
  • Sentry canary and full-fleet soak pass.
  • Signed-APT rollout and rollback are proven.

A broader one-process-per-server Tokio redesign remains optional until these gates pass and an explicit failure-domain decision is recorded.

Non-goals

  • Do not remove or simplify velnor-actions-fixture content.
  • Do not hide the problem with a repository-local workflow workaround.
  • Do not broad-prune Docker or delete data without ownership proof.
  • Do not trade crash durability or job isolation for lower idle CPU.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions