You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On 2026-08-26, Sentry consumed roughly 600–700% CPU while docker ps showed no containers and no user jobs running. The host had about 9–10 active velnor-runner daemon processes, each daemon main thread using approximately 50–84% CPU. Memory was normal (~9.5 GiB / 62.6 GiB).
This is control-plane work, not job execution.
Observed Sentry symptoms:
repeated registration lost messages;
failed JIT registration cleanup/delete;
repeated JIT reconfiguration attempts;
broker session HTTP 409 conflicts;
broker poll HTTP 404 failures;
stale GitHub registrations and timed-out unassigned jobs;
many idle slot and waiter/job processes despite no Docker containers;
doctor reporting degraded/failed fleet state while daemon health could still report ready.
Root cause
The current topology multiplies idle control work:
ready slot
-> separate slot process
-> separate waiter/job process
-> separate broker session and poll loop
-> registration/session drift
-> JIT cleanup/re-registration churn
-> repeated controller reconciliation and durable journal writes
-> high CPU with zero Docker jobs
The strongest deterministic hot path is repeated reconciliation and journal write amplification. Every controller reconciles every two seconds, repeatedly applies proof/routing/capacity observations, and performs durable SQLite work. The per-slot waiter/job processes multiply this work and each owns independent broker/session recovery.
crates/velnor-runner/src/runner.rs:3189-3281 — broker retry state is local to each runner.
crates/velnor-runner/src/protocol.rs:1841-1867 — broker 404 enters the error path.
crates/velnor-control/src/journal.rs:932-999 — each accepted event performs durable SQLite work.
Architecture decision
Use Tokio to consolidate idle control-plane concurrency, but do not immediately collapse the whole server into one address space.
The current normative Node Architecture v2 requires guardian, per-scope controller, per-ready-slot process, and transient job worker (docs/vision.md:82-91, docs/roadmap.md:31-35, docs/mission.md:50-55). This is a deliberate failure-containment boundary.
Target architecture:
guardian
└── controller per scope
├── bounded async broker/session manager
├── event-driven reconciliation
├── isolated slot lifecycle
└── transient job process → Docker or microVM execution
A single node-wide process is a later experiment only if it proves equivalent containment for panic, deadlock, protocol stall, CPU spin, OOM, and restart. One scope’s failure must never take down unrelated scopes or active jobs.
Implementation plan
Phase 0 — establish attribution and budgets
Add reconcile cycle count, duration p50/p95/p99, and overlap detection.
Add events/sec and accepted/rejected/no-op counts by event kind.
Problem
On 2026-08-26, Sentry consumed roughly 600–700% CPU while
docker psshowed no containers and no user jobs running. The host had about 9–10 activevelnor-runner daemonprocesses, each daemon main thread using approximately 50–84% CPU. Memory was normal (~9.5 GiB / 62.6 GiB).This is control-plane work, not job execution.
Observed Sentry symptoms:
registration lostmessages;Root cause
The current topology multiplies idle control work:
The strongest deterministic hot path is repeated reconciliation and journal write amplification. Every controller reconciles every two seconds, repeatedly applies proof/routing/capacity observations, and performs durable SQLite work. The per-slot waiter/job processes multiply this work and each owns independent broker/session recovery.
Relevant code:
crates/velnor-runner/src/node/controller.rs:232-275— unconditional two-second reconciliation.crates/velnor-runner/src/node/controller.rs:339-425— per-cycle proof, registration, executor, session, routing, and journal work.crates/velnor-runner/src/node/controller.rs:803-842— waiter for every ready slot.crates/velnor-runner/src/node/controller.rs:1002-1045— waiter as separatevelnor-runner jobprocess.crates/velnor-runner/src/node/job.rs:48-70— waiter runs the full broker runner lifecycle.crates/velnor-runner/src/runner.rs:1208-1215— one OS process per slot, including surge capacity.crates/velnor-runner/src/runner.rs:1411-1495— lost registration triggers cleanup, fresh JIT configuration, and another slot cycle.crates/velnor-runner/src/runner.rs:2666-2746— independent idle broker polling loop.crates/velnor-runner/src/runner.rs:3189-3281— broker retry state is local to each runner.crates/velnor-runner/src/protocol.rs:1841-1867— broker 404 enters the error path.crates/velnor-control/src/journal.rs:932-999— each accepted event performs durable SQLite work.Architecture decision
Use Tokio to consolidate idle control-plane concurrency, but do not immediately collapse the whole server into one address space.
The current normative Node Architecture v2 requires guardian, per-scope controller, per-ready-slot process, and transient job worker (
docs/vision.md:82-91,docs/roadmap.md:31-35,docs/mission.md:50-55). This is a deliberate failure-containment boundary.Target architecture:
A single node-wide process is a later experiment only if it proves equivalent containment for panic, deadlock, protocol stall, CPU spin, OOM, and restart. One scope’s failure must never take down unrelated scopes or active jobs.
Implementation plan
Phase 0 — establish attribution and budgets
jobs,idle_slots, request rates, event rates, and resource status to doctor/health.Phase 1 — remove journal write amplification
PermitReserved,ExecutorProven,SessionLive, routing, and dependency observations no-op when unchanged.apply_many()transaction.synchronous=FULLand crash durability.Phase 2 — eliminate idle waiter processes
velnor-runner jobwaiter.Phase 3 — centralize broker and registration recovery
Healthy → Missing/Conflict → Quarantined → Recreate → Registered.202,401,403,404,409, timeout/reset, malformed response, rate limit, and registration disappearance.Phase 4 — bound reconciliation
READY=1and health distinguish control-cycle liveness from schedulable capacity.jobs=0 + high CPU.Phase 5 — tests and failure isolation
401,403,404,409, timeout, malformed payload, rate limit, delete failure, and JIT failure.cargo nextest run.Phase 6 — fixture and performance proof
velnor-actions-fixtureunchanged.Initial gates; calibrate only from the recorded baseline:
Phase 7 — canary and release
Ready definition
Close this issue only when all are true:
A broader one-process-per-server Tokio redesign remains optional until these gates pass and an explicit failure-domain decision is recorded.
Non-goals
velnor-actions-fixturecontent.