Warm-respawn a supervisor unit whose host was lost - #2009
Merged
ppXD merged 1 commit intoSep 23, 2026
Merged
Conversation
#1994 made a lost host's agent continuable, but only for a workflow agent.run node: the observer checkpoints the live transcript, the reconciler's abandon keeps the reference, and the node's own respawn consumes it. A supervisor-spawned unit was excluded — its envelope never opted in, so its retry restarted the conversation from zero even though the machinery to continue it was already there. The staging loop that every spawn wave, retry and resolve passes through now opts a unit in only when a later retry could consume the checkpoint: the unit carries a subtask id (resolver units carry none, and the retry lookup matches on it), and SupervisorBounds.CanRespawnAfterWave says the run can still afford to respawn it. The subtask-attempt lookup falls back to a TERMINAL row's checkpoint columns when the attempt left no result to read, which is the shape an abandon leaves. A Running row's checkpoint belongs to a session that may still be live, so it is never resumed. For "terminal and still holding a checkpoint" to mean an abandon, a deliberate cancel now releases the columns as completion does; only the reconciler's abandon and spool recovery keep them. The resume fold tells the two kinds of restored conversation apart: a capture finished with its tree intact, a checkpoint did not, so only the second carries the lost-host block and its provenance. Its tree sentence reads the resumed attempt's own pushed branch, not the effective clone ref, so a dependent unit is never told that its producer's branch is its own published work. The same change stops a producer's handoff ref from suppressing the honest-redo line on the ordinary path when the prior attempt pushed nothing itself. The billing stays honest: the graded roster keys by subtask and takes the latest, so one host loss spends one graded attempt, not two.
ppXD
force-pushed
the
feat/warm-respawn-a-supervisor-unit-whose-host-was-lost
branch
from
September 23, 2026 07:19
2259b18 to
df9d1fb
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
RealSupervisorActionExecutor.Spawn.cs:824) callsCheckpointsSessionTranscript(:174). That requires a subtask id, which resolver units never carry and the retry lookup matches on. It also requiresSupervisorBounds.CanRespawnAfterWave(SupervisorBounds.cs:109), which reads the same spawn capPostDecisionenforces, strictly below it, for the whole wave.FindResumableSubtaskAttemptAsyncfalls back to a terminal row's checkpoint columns when the attempt left no result (AgentRunService.cs:981, guarded at:999). A Running row's checkpoint belongs to a session that may still be live, so it is never resumed.ApplyResumeRecord(Spawn.cs:561) stages the checkpoint by reference with its provenance and the lost-host block. A captured transcript gets none of that, and an unreadable checkpoint still starts cold throughAgentRunExecutor.ResolveCheckpointTranscriptAsync.AgentRunService.cs:788). A terminal row that still holds a checkpoint therefore means an abandon, and a cancelled unit no longer keeps its transcript artifact for good. The reconciler's spool recovery (AgentRunReconcilerService.cs:934) still keeps the checkpoint on purpose: a process that died on a live host is an abandon-class ending, and its checkpoint is the retry's best conversation source.workspaceRef: priorAttemptStaging.Ref,Spawn.cs:453) instead of the effective clone ref, so a dependent unit is never told that its producer's branch is its own work.InfraExitReason. The agent.run lane is unaffected: it never respawns a cancelled run (AgentCodeNode.cs:379).AgentRunReconcilerService.cs:378,387,396), including spool recovery.Test plan
SupervisorBoundsTestscovers both predicate arms, plus headroom agreeing withPostDecisionat an explicit cap and at the lane-default fallback.SupervisorDependencyStagingTestscovers the fold and asserts exactly which tree sentence is emitted: published branch, honest-redo, or none when there is no repository.SupervisorDependencyGateTestscovers a respawn replacing the abandoned attempt in the graded roster.SupervisorRetryWorldStateFlowTests:CancelRunningAsyncclears both columns and leaves the retry coldSupervisorResolveFlowTests: a resolver unit never opts in.