Retry a lost host's agent warm instead of cold - #1994
Merged
ppXD merged 1 commit intoSep 19, 2026
Merged
Conversation
ppXD
force-pushed
the
feat/resume-a-lost-hosts-run-from-a-durable-checkpoint
branch
3 times, most recently
from
September 19, 2026 06:00
24a601b to
dce89a9
Compare
ppXD
force-pushed
the
feat/resume-a-lost-hosts-run-from-a-durable-checkpoint
branch
2 times, most recently
from
September 19, 2026 10:20
8f115d7 to
9c0d893
Compare
An agent node whose host dies is already retried today: the reconciler abandons the run (Failed, no result), the agent.run node grades that as a retryable failure, and the node's retry policy stages a fresh agent. The retry is COLD. The conversation the dead machine was holding is simply gone, because the resumable CLI transcript was captured only at run END, out of the launching host's own spool — which no other worker can read and which dies with the host. The observer's drain tick now makes that transcript durable while the run is alive. It reads the cwd the CLI actually ran in, since Claude keys its session file on that path and the primary repository's directory is not it for a multi-repo workspace and does not exist at all for a repo-less one. It stores the complete prefix, cut at the last newline, because the CLI is still appending and a half-written trailing line is a corrupt session rather than a shorter one. It verifies through /proc/self/fd that the file it opened is the one the clamp resolved, since the walk and the open are two syscalls with a live writer between them. And it runs off the tick, behind a claimed cadence window, so the drain loop carrying stdout, the exit marker and the progress lease never waits on an upload. The abandon leaves the reference on the row — deliberately, where the executor's own terminal write clears it — the completion notifier projects it across the wait boundary, and the node's existing respawn stamps it onto the fresh task as a transcript ref plus the session id. Only the conversation survives: the dead host's working tree does not, so the goal carries the lost-host block, the run records that it resumed from a checkpoint, and a column names the attempt it took over from. Checkpointing is an opt-in the agent node sets from the engine's own answer about its retry policy, so a run whose failure nobody can retry — a benchmark cell, a review child, a supervisor unit, a single-attempt node — pays for nothing. Checkpoints are their own retention class with a two-hour floor: a run writes one a minute, each supersedes the last, and the growth gate that makes them worth taking is exactly what stops the content-addressed store deduplicating them. Supervisor units are out of scope; their respawn can read the same columns through AgentRetryCauses later.
ppXD
force-pushed
the
feat/resume-a-lost-hosts-run-from-a-durable-checkpoint
branch
from
September 19, 2026 10:41
9c0d893 to
19a24a0
Compare
7 tasks done
ppXD
added a commit
that referenced
this pull request
Sep 23, 2026
#1994 made a lost host's agent continuable, but only for a workflow agent.run node: the observer checkpoints the live transcript, the reconciler's abandon keeps the reference, and the node's own respawn consumes it. A supervisor-spawned unit was excluded — its envelope never opted in, so its retry restarted the conversation from zero even though the machinery to continue it was already there. The staging loop that every spawn wave, retry and resolve passes through now opts a unit in only when a later retry could consume the checkpoint: the unit carries a subtask id (resolver units carry none, and the retry lookup matches on it), and SupervisorBounds.CanRespawnAfterWave says the run can still afford to respawn it. The subtask-attempt lookup falls back to a TERMINAL row's checkpoint columns when the attempt left no result to read, which is the shape an abandon leaves. A Running row's checkpoint belongs to a session that may still be live, so it is never resumed. For "terminal and still holding a checkpoint" to mean an abandon, a deliberate cancel now releases the columns as completion does; only the reconciler's abandon and spool recovery keep them. The resume fold tells the two kinds of restored conversation apart: a capture finished with its tree intact, a checkpoint did not, so only the second carries the lost-host block and its provenance. Its tree sentence reads the resumed attempt's own pushed branch, not the effective clone ref, so a dependent unit is never told that its producer's branch is its own published work. The same change stops a producer's handoff ref from suppressing the honest-redo line on the ordinary path when the prior attempt pushed nothing itself. The billing stays honest: the graded roster keys by subtask and takes the latest, so one host loss spends one graded attempt, not two.
ppXD
added a commit
that referenced
this pull request
Sep 23, 2026
#1994 made a lost host's agent continuable, but only for a workflow agent.run node: the observer checkpoints the live transcript, the reconciler's abandon keeps the reference, and the node's own respawn consumes it. A supervisor-spawned unit was excluded — its envelope never opted in, so its retry restarted the conversation from zero even though the machinery to continue it was already there. The staging loop that every spawn wave, retry and resolve passes through now opts a unit in only when a later retry could consume the checkpoint: the unit carries a subtask id (resolver units carry none, and the retry lookup matches on it), and SupervisorBounds.CanRespawnAfterWave says the run can still afford to respawn it. The subtask-attempt lookup falls back to a TERMINAL row's checkpoint columns when the attempt left no result to read, which is the shape an abandon leaves. A Running row's checkpoint belongs to a session that may still be live, so it is never resumed. For "terminal and still holding a checkpoint" to mean an abandon, a deliberate cancel now releases the columns as completion does; only the reconciler's abandon and spool recovery keep them. The resume fold tells the two kinds of restored conversation apart: a capture finished with its tree intact, a checkpoint did not, so only the second carries the lost-host block and its provenance. Its tree sentence reads the resumed attempt's own pushed branch, not the effective clone ref, so a dependent unit is never told that its producer's branch is its own published work. The same change stops a producer's handoff ref from suppressing the honest-redo line on the ordinary path when the prior attempt pushed nothing itself. The billing stays honest: the graded roster keys by subtask and takes the latest, so one host loss spends one graded attempt, not two.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
What production does today for a host-loss abandon of an agent node: it RETRIES — cold. The reconciler terminalizes the run (
Failed, the abandoned error,ResultJsonnull),AgentCodeNodegrades it retryable (it is not one of the four deterministic carve-outs), and the node's retry policy stages a fresh agent run. Pinned byA_host_loss_abandon_of_an_agent_node_buys_a_second_attemptbefore anything was built on it. This PR does not buy that attempt — it makes it warm.Produce. The observer's existing drain tick checkpoints the live session transcript. Keyed on the cwd the CLI actually ran in (Claude keys its session file on that path; the primary repo's directory is not it for a multi-repo workspace and does not exist for a repo-less one). Stores the complete prefix cut at the last newline — the CLI is still appending, and a half-written line is a corrupt session, not a shorter one. Verifies through
/proc/self/fdthat the opened file is the one the clamp resolved, because the walk and the open are two syscalls with a live writer between them. Runs off the tick, in a DI scope of its own, behind a cadence window the executor owns — the tick is using the executor's scopedDbContextevery 250 ms, so a checkpoint sharing it would put two operations on one EF context and kill a healthy run for a best-effort aid. It carries a deadline of its own rather than the observer's token: that token is cancelled by exactly the terminations a warm retry follows (wall-clock timeout, no-progress stall, operator cancel), so threading it would abort the last checkpoint on the runs whose conversation is most worth keeping. The round waits for that checkpoint before completing (the stamp is fenced on the run still being Running), in afinallyso the paths a retry follows do not skip it.Two budgets, because they bound two different things:
SessionCheckpointUploadBudget(30 s) is how long one checkpoint may take — sized for the 32 MiB capture cap, so a throttled or cross-region destination does not cancel every checkpoint of exactly the long conversations this protects — andSessionCheckpointDrainBudget(5 s) is how long a finished run's terminal write may be deferred by one. Past the drain bound the run lands and a late stamp is refused by the fence.Degrade. A checkpoint ref that cannot be read (blob reaped, destination unreachable) makes the retry run cold, not fail — the task now says which kind of ref it carries, and only a captured transcript still fails closed. Spending the retry on the one fault that says nothing about the work would defeat the point.
Carry. The reconciler's abandon deliberately keeps the checkpoint columns (the executor's own terminal write clears them); the completion notifier projects them off the row — an abandon has no result, which is exactly why the existing captured-transcript keys are null for this population.
Consume.
AgentCodeNode.ApplyRespawnResumeHintstamps the fresh task with the session id, the checkpoint as a ref,ResumedFromCheckpointAt,ResumedFromAgentRunId, and the lost-host continuity block. The working tree died with the host, so the agent is told — even when a branch was pinned, since the unpublished remainder is still gone.Opt-in.
AgentTask.CheckpointSessionTranscript, set byAgentCodeNodefromNodeRunContext.RetriesOnFailure— the engine's own answer about this node's retry policy. A benchmark cell, a review child, a supervisor unit or a single-attempt node pays for nothing.Retention. Checkpoints get their own class with a two-hour floor (a run writes one a minute, each supersedes the last, and the growth gate that makes them worth taking is what stops the store deduplicating them), and the clean terminal write releases the survivor so the reaper can collect it.
Migration
0234: two nullable checkpoint columns +resumed_from_agent_run_id, plus a partial index on the checkpoint artifact id — the reference oracle probes it, and0144's own rule is that an unindexed site turns eachEXISTSinto a sequential scan ofagent_run.Retention cost, stated because it is permanent. The abandon KEEPS the checkpoint column (that is the whole mechanism), which keeps the artifact
Referenced— andReferencedis terminal in the retention ledger, so one full transcript copy is kept for ever per host loss, including when the retry ran and finished. The two-hour class does not collect it; that class only shortens the age floor for checkpoints nothing references. Clearing the column once the successor is staged is not available either: the successor's ref lives in its owntask_json, which the oracle does not probe, so the artifact would become collectable in the window before that successor launches. A jsonb reference site is the fix and belongs with the retention plane.Removed from the first cut, after the enumeration showed no population for it: the reconciler-bought standalone continuation (
ResumeFromCheckpointAsync, its create seam,AdmitCheckpointResumeAsync,MaxResumeAttempts, the enqueue and its timeline note). Every operator launch is workflow-embedded; the onlyworkflow_run_id IS NULLproducers are the benchmark lanes, whose protocol is one attempt per cell. A correct-but-dormant mechanism is speculative code.Follow-up (not in scope): supervisor units. Their respawn goes through
AgentRetryCauses, not the node, and can read the same two columns later.Test plan
AgentSessionTranscriptCheckpointerTests(15) — the growth watermark's atomicity across a revise round, the cadence gate and the checkpoint budget, the opened-handle identity check, growth, byte cap + its under-cap pair, torn trailing line, no-complete-line, growth inside the torn tail, a new round's smaller transcript, absent file, lost fence.AgentCodeNodeTests(4 new) — host-loss restore, cold-when-no-checkpoint, and the opt-in Theory. Plus the oracle and retention pins.AgentRunSessionCheckpointFlowTests(12) — a slow destination still gets its checkpoint stored while a wedged one cannot defer the terminal write; an unreadable checkpoint ref degrades to a cold start while an unreadable captured ref still fails closed; a wedged store cannot defer the terminal write; a streaming agent survives a checkpoint that is still uploading; the last checkpoint lands even when the agent exits immediately; the REAL executor + drain tick + checkpointer + artifact store checkpoint a live agent for a repo-less and an authored-cwd run; a run that did not opt in writes nothing; a landed run's checkpoint readsUnreferenced; the fenced stamp lands, refuses a superseded fence and never nulls a session id; DI wiring.AgentNodeFlowTests(3 new) — the retryability pin, the warm restore end to end, and the cold retry that claims nothing.RerunFromNodeAgentFlowTests.An_in_flight_sibling_at_one_hop_cannot_mask_the_resumable_attempt(a run now stamps its session id at its first checkpoint, so an in-flight sibling is a real masking hazard; the hop is now newest-first and scans).RerunFromNodeAgentFlowTests8/8. Rebased ontoabd38280b(Reclaim log-stream bytes nothing cites any more #1993): both PRs'ReferenceSitesentries kept, 0234 unchanged,MigrationDiscoveryTestsgreen.RealModelSessionResumeE2ETests.A_real_claude_agent_recalls_a_codeword_from_a_mid_run_checkpoint— informational; its doc states exactly what it proves (a mid-run transcript located through the production locate + clamp is bytes a realclauderesumes from) and that it covers none of this PR's plumbing.Mutations verified red
A_new_rounds_watermark_is_never_paired_with_the_old_rounds_byte_countLandedmade a no-opA_landed_checkpoint_is_what_the_next_attempt_must_grow_pastA_slow_destination_still_gets_its_checkpoint_stored— nothing was ever stored against an 8 s destinationA_wedged_store_cannot_defer_the_terminal_write_past_the_checkpoint_budget— the round took 21.2sAn_unreadable_transcript_ref_degrades_only_when_it_is_a_checkpointDbContext)A_streaming_agent_survives_a_checkpoint_that_is_still_uploading— with the literal "A second operation was started on this context instance"The_last_checkpoint_lands_even_when_the_agent_exits_immediatelyOpenedTheResolvedFilereplaced withreturn trueA_handle_that_is_not_the_resolved_file_is_refusedThe_checkpoint_cadence_reopens_only_after_its_intervalA_run_that_did_not_opt_in_writes_no_checkpoint_at_all