Skip to content

Warm-respawn a supervisor unit whose host was lost - #2009

Merged
ppXD merged 1 commit into
mainfrom
feat/warm-respawn-a-supervisor-unit-whose-host-was-lost
Sep 23, 2026
Merged

ppXD merged 1 commit into
mainfrom
feat/warm-respawn-a-supervisor-unit-whose-host-was-lost

Conversation

@ppXD

@ppXD ppXD commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Supervisor units now opt into the session checkpoint from Retry a lost host's agent warm instead of cold #1994, but only when a later retry could consume it. The staging loop that every spawn wave, retry and resolve passes through (RealSupervisorActionExecutor.Spawn.cs:824) calls CheckpointsSessionTranscript (:174). That requires a subtask id, which resolver units never carry and the retry lookup matches on. It also requires SupervisorBounds.CanRespawnAfterWave (SupervisorBounds.cs:109), which reads the same spawn cap PostDecision enforces, strictly below it, for the whole wave.
  • A retried unit now resumes from the checkpoint when its host was lost. FindResumableSubtaskAttemptAsync falls back to a terminal row's checkpoint columns when the attempt left no result (AgentRunService.cs:981, guarded at :999). A Running row's checkpoint belongs to a session that may still be live, so it is never resumed. ApplyResumeRecord (Spawn.cs:561) stages the checkpoint by reference with its provenance and the lost-host block. A captured transcript gets none of that, and an unreadable checkpoint still starts cold through AgentRunExecutor.ResolveCheckpointTranscriptAsync.
  • Every claim the retry makes now describes the right attempt:
    • A deliberate cancel releases the checkpoint columns, as completion does (AgentRunService.cs:788). A terminal row that still holds a checkpoint therefore means an abandon, and a cancelled unit no longer keeps its transcript artifact for good. The reconciler's spool recovery (AgentRunReconcilerService.cs:934) still keeps the checkpoint on purpose: a process that died on a live host is an abandon-class ending, and its checkpoint is the retry's best conversation source.
    • The tree sentence reads the resumed attempt's own pushed branch (workspaceRef: priorAttemptStaging.Ref, Spawn.cs:453) instead of the effective clone ref, so a dependent unit is never told that its producer's branch is its own work.
    • The same change affects the ordinary path too. A producer's handoff ref no longer suppresses the honest-redo line when the prior attempt pushed nothing itself.
  • Grading is unchanged. The graded roster keeps only the latest attempt per subtask, so one host loss is billed as one attempt. A host-loss abandon still grades as an ordinary failure, not infra, because it has no result and so no InfraExitReason. The agent.run lane is unaffected: it never respawns a cancelled run (AgentCodeNode.cs:379).
  • Known and out of scope:
    • The "machine was lost" wording from Retry a lost host's agent warm instead of cold #1994 is used for every abandon cause (AgentRunReconcilerService.cs:378,387,396), including spool recovery.
    • The existing lookup lets an older resumable attempt outrank a newer attempt's pushed-but-uncaptured work.

Test plan

  • Unit: SupervisorBoundsTests covers both predicate arms, plus headroom agreeing with PostDecision at an explicit cap and at the lane-default fallback. SupervisorDependencyStagingTests covers the fold and asserts exactly which tree sentence is emitted: published branch, honest-redo, or none when there is no repository. SupervisorDependencyGateTests covers a respawn replacing the abandoned attempt in the graded roster.
  • Integration, SupervisorRetryWorldStateFlowTests:
    • host loss with a checkpoint, abandoned by the real reconciler, resumes warm
    • host loss without a checkpoint resumes cold and claims nothing
    • a deliberate cancel through the real CancelRunningAsync clears both columns and leaves the retry cold
    • a still-Running prior attempt is never resumed
    • dependent units: lost before pushing (honest-redo, never the producer's branch), captured transcript without a push (honest-redo), and pushed its own branch (that branch is named)
    • the opt-in at the staging seam for a retry and for a two-agent wave
  • Integration, SupervisorResolveFlowTests: a resolver unit never opts in.
  • Each new guard fails its test under a named mutation: effective ref restored, own ref dropped, cancel clearing dropped, status guard dropped, subtask-id gate dropped, bare-session-id guard (both variants), wave size replaced by 1.
  • Full unit suite: 10931 passed, 1 skipped.
  • Supervisor, reconciler, executor and agent-run cancel integration classes: 724 passed, 16 skipped (real-model, no credentials).
  • E2E: no new arm, since the warm path needs a host that never comes back. The failure-to-retry arm's doc no longer claims coverage it does not assert.

#1994 made a lost host's agent continuable, but only for a workflow
agent.run node: the observer checkpoints the live transcript, the
reconciler's abandon keeps the reference, and the node's own respawn
consumes it. A supervisor-spawned unit was excluded — its envelope
never opted in, so its retry restarted the conversation from zero even
though the machinery to continue it was already there.

The staging loop that every spawn wave, retry and resolve passes
through now opts a unit in only when a later retry could consume the
checkpoint: the unit carries a subtask id (resolver units carry none,
and the retry lookup matches on it), and
SupervisorBounds.CanRespawnAfterWave says the run can still afford to
respawn it.

The subtask-attempt lookup falls back to a TERMINAL row's checkpoint
columns when the attempt left no result to read, which is the shape an
abandon leaves. A Running row's checkpoint belongs to a session that
may still be live, so it is never resumed. For "terminal and still
holding a checkpoint" to mean an abandon, a deliberate cancel now
releases the columns as completion does; only the reconciler's abandon
and spool recovery keep them.

The resume fold tells the two kinds of restored conversation apart: a
capture finished with its tree intact, a checkpoint did not, so only
the second carries the lost-host block and its provenance. Its tree
sentence reads the resumed attempt's own pushed branch, not the
effective clone ref, so a dependent unit is never told that its
producer's branch is its own published work. The same change stops a
producer's handoff ref from suppressing the honest-redo line on the
ordinary path when the prior attempt pushed nothing itself.

The billing stays honest: the graded roster keys by subtask and takes
the latest, so one host loss spends one graded attempt, not two.
@ppXD
ppXD force-pushed the feat/warm-respawn-a-supervisor-unit-whose-host-was-lost branch from 2259b18 to df9d1fb Compare September 23, 2026 07:19
@ppXD
ppXD merged commit 475c722 into main Sep 23, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant