Skip to content

Retry a lost host's agent warm instead of cold - #1994

Merged
ppXD merged 1 commit into
mainfrom
feat/resume-a-lost-hosts-run-from-a-durable-checkpoint
Sep 19, 2026
Merged

ppXD merged 1 commit into
mainfrom
feat/resume-a-lost-hosts-run-from-a-durable-checkpoint

Conversation

@ppXD

@ppXD ppXD commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Summary

What production does today for a host-loss abandon of an agent node: it RETRIES — cold. The reconciler terminalizes the run (Failed, the abandoned error, ResultJson null), AgentCodeNode grades it retryable (it is not one of the four deterministic carve-outs), and the node's retry policy stages a fresh agent run. Pinned by A_host_loss_abandon_of_an_agent_node_buys_a_second_attempt before anything was built on it. This PR does not buy that attempt — it makes it warm.

  • Produce. The observer's existing drain tick checkpoints the live session transcript. Keyed on the cwd the CLI actually ran in (Claude keys its session file on that path; the primary repo's directory is not it for a multi-repo workspace and does not exist for a repo-less one). Stores the complete prefix cut at the last newline — the CLI is still appending, and a half-written line is a corrupt session, not a shorter one. Verifies through /proc/self/fd that the opened file is the one the clamp resolved, because the walk and the open are two syscalls with a live writer between them. Runs off the tick, in a DI scope of its own, behind a cadence window the executor owns — the tick is using the executor's scoped DbContext every 250 ms, so a checkpoint sharing it would put two operations on one EF context and kill a healthy run for a best-effort aid. It carries a deadline of its own rather than the observer's token: that token is cancelled by exactly the terminations a warm retry follows (wall-clock timeout, no-progress stall, operator cancel), so threading it would abort the last checkpoint on the runs whose conversation is most worth keeping. The round waits for that checkpoint before completing (the stamp is fenced on the run still being Running), in a finally so the paths a retry follows do not skip it.

    Two budgets, because they bound two different things: SessionCheckpointUploadBudget (30 s) is how long one checkpoint may take — sized for the 32 MiB capture cap, so a throttled or cross-region destination does not cancel every checkpoint of exactly the long conversations this protects — and SessionCheckpointDrainBudget (5 s) is how long a finished run's terminal write may be deferred by one. Past the drain bound the run lands and a late stamp is refused by the fence.

  • Degrade. A checkpoint ref that cannot be read (blob reaped, destination unreachable) makes the retry run cold, not fail — the task now says which kind of ref it carries, and only a captured transcript still fails closed. Spending the retry on the one fault that says nothing about the work would defeat the point.

  • Carry. The reconciler's abandon deliberately keeps the checkpoint columns (the executor's own terminal write clears them); the completion notifier projects them off the row — an abandon has no result, which is exactly why the existing captured-transcript keys are null for this population.

  • Consume. AgentCodeNode.ApplyRespawnResumeHint stamps the fresh task with the session id, the checkpoint as a ref, ResumedFromCheckpointAt, ResumedFromAgentRunId, and the lost-host continuity block. The working tree died with the host, so the agent is told — even when a branch was pinned, since the unpublished remainder is still gone.

  • Opt-in. AgentTask.CheckpointSessionTranscript, set by AgentCodeNode from NodeRunContext.RetriesOnFailure — the engine's own answer about this node's retry policy. A benchmark cell, a review child, a supervisor unit or a single-attempt node pays for nothing.

  • Retention. Checkpoints get their own class with a two-hour floor (a run writes one a minute, each supersedes the last, and the growth gate that makes them worth taking is what stops the store deduplicating them), and the clean terminal write releases the survivor so the reaper can collect it.

Migration 0234: two nullable checkpoint columns + resumed_from_agent_run_id, plus a partial index on the checkpoint artifact id — the reference oracle probes it, and 0144's own rule is that an unindexed site turns each EXISTS into a sequential scan of agent_run.

Retention cost, stated because it is permanent. The abandon KEEPS the checkpoint column (that is the whole mechanism), which keeps the artifact Referenced — and Referenced is terminal in the retention ledger, so one full transcript copy is kept for ever per host loss, including when the retry ran and finished. The two-hour class does not collect it; that class only shortens the age floor for checkpoints nothing references. Clearing the column once the successor is staged is not available either: the successor's ref lives in its own task_json, which the oracle does not probe, so the artifact would become collectable in the window before that successor launches. A jsonb reference site is the fix and belongs with the retention plane.

Removed from the first cut, after the enumeration showed no population for it: the reconciler-bought standalone continuation (ResumeFromCheckpointAsync, its create seam, AdmitCheckpointResumeAsync, MaxResumeAttempts, the enqueue and its timeline note). Every operator launch is workflow-embedded; the only workflow_run_id IS NULL producers are the benchmark lanes, whose protocol is one attempt per cell. A correct-but-dormant mechanism is speculative code.

Follow-up (not in scope): supervisor units. Their respawn goes through AgentRetryCauses, not the node, and can read the same two columns later.

Test plan

  • Unit (19, both budget literals pinned): AgentSessionTranscriptCheckpointerTests (15) — the growth watermark's atomicity across a revise round, the cadence gate and the checkpoint budget, the opened-handle identity check, growth, byte cap + its under-cap pair, torn trailing line, no-complete-line, growth inside the torn tail, a new round's smaller transcript, absent file, lost fence. AgentCodeNodeTests (4 new) — host-loss restore, cold-when-no-checkpoint, and the opt-in Theory. Plus the oracle and retention pins.
  • Integration against Postgres: AgentRunSessionCheckpointFlowTests (12) — a slow destination still gets its checkpoint stored while a wedged one cannot defer the terminal write; an unreadable checkpoint ref degrades to a cold start while an unreadable captured ref still fails closed; a wedged store cannot defer the terminal write; a streaming agent survives a checkpoint that is still uploading; the last checkpoint lands even when the agent exits immediately; the REAL executor + drain tick + checkpointer + artifact store checkpoint a live agent for a repo-less and an authored-cwd run; a run that did not opt in writes nothing; a landed run's checkpoint reads Unreferenced; the fenced stamp lands, refuses a superseded fence and never nulls a session id; DI wiring. AgentNodeFlowTests (3 new) — the retryability pin, the warm restore end to end, and the cold retry that claims nothing.
  • E2E: RerunFromNodeAgentFlowTests.An_in_flight_sibling_at_one_hop_cannot_mask_the_resumable_attempt (a run now stamps its session id at its first checkpoint, so an in-flight sibling is a real masking hazard; the hop is now newest-first and scans).
  • Full unit suite 10850/10851 (1 pre-existing skip); touched integration classes 533/533 (227 re-run after the budget split); RerunFromNodeAgentFlowTests 8/8. Rebased onto abd38280b (Reclaim log-stream bytes nothing cites any more #1993): both PRs' ReferenceSites entries kept, 0234 unchanged, MigrationDiscoveryTests green.
  • Real-model lane: RealModelSessionResumeE2ETests.A_real_claude_agent_recalls_a_codeword_from_a_mid_run_checkpoint — informational; its doc states exactly what it proves (a mid-run transcript located through the production locate + clamp is bytes a real claude resumes from) and that it covers none of this PR's plumbing.

Mutations verified red

Mutation Reds
growth watermark's path advanced on dispatch while its byte count advanced on landing A_new_rounds_watermark_is_never_paired_with_the_old_rounds_byte_count
Landed made a no-op that test and A_landed_checkpoint_is_what_the_next_attempt_must_grow_past
upload budget collapsed into the drain's 5 s A_slow_destination_still_gets_its_checkpoint_stored — nothing was ever stored against an 8 s destination
drain bound widened to the upload's 30 s A_wedged_store_cannot_defer_the_terminal_write_past_the_checkpoint_budget — the round took 21.2s
both transcript refs fail closed / both degrade the two arms of An_unreadable_transcript_ref_degrades_only_when_it_is_a_checkpoint
checkpoint resolved from the executor's own scope (shared DbContext) A_streaming_agent_survives_a_checkpoint_that_is_still_uploading — with the literal "A second operation was started on this context instance"
the round does not wait for its in-flight checkpoint The_last_checkpoint_lands_even_when_the_agent_exits_immediately
OpenedTheResolvedFile replaced with return true A_handle_that_is_not_the_resolved_file_is_refused
the cadence gate always open The_checkpoint_cadence_reopens_only_after_its_interval
host-loss abandon graded non-retryable the retryability pin + 4 existing respawn tests
node's checkpoint consumer removed warm-restore test (unit + integration)
notifier stops carrying the checkpoint across the wait warm-restore integration test
lost-host block stamped unconditionally cold-retry test + 3 existing honesty arms
checkpoint opt-in predicate removed the opt-in Theory
tick keyed off the primary repo directory both real-tick arms + the release test
tick built without the opt-in gate A_run_that_did_not_opt_in_writes_no_checkpoint_at_all
no truncation at the last newline 3 checkpointer unit tests
terminal write keeps the checkpoint reference 3 integration tests
stamp loses its fence / overwrites the session id the stamp test
hop takes the first candidate instead of scanning the masking E2E
growth / cadence / byte cap / path key removed the four corresponding unit tests

@ppXD
ppXD force-pushed the feat/resume-a-lost-hosts-run-from-a-durable-checkpoint branch 3 times, most recently from 24a601b to dce89a9 Compare September 19, 2026 06:00
@ppXD ppXD changed the title Continue a lost host's run from a durable session checkpoint Retry a lost host's agent warm instead of cold Sep 19, 2026
@ppXD
ppXD force-pushed the feat/resume-a-lost-hosts-run-from-a-durable-checkpoint branch 2 times, most recently from 8f115d7 to 9c0d893 Compare September 19, 2026 10:20
An agent node whose host dies is already retried today: the reconciler
abandons the run (Failed, no result), the agent.run node grades that as a
retryable failure, and the node's retry policy stages a fresh agent. The
retry is COLD. The conversation the dead machine was holding is simply
gone, because the resumable CLI transcript was captured only at run END,
out of the launching host's own spool — which no other worker can read and
which dies with the host.

The observer's drain tick now makes that transcript durable while the run
is alive. It reads the cwd the CLI actually ran in, since Claude keys its
session file on that path and the primary repository's directory is not it
for a multi-repo workspace and does not exist at all for a repo-less one.
It stores the complete prefix, cut at the last newline, because the CLI is
still appending and a half-written trailing line is a corrupt session
rather than a shorter one. It verifies through /proc/self/fd that the file
it opened is the one the clamp resolved, since the walk and the open are
two syscalls with a live writer between them. And it runs off the tick,
behind a claimed cadence window, so the drain loop carrying stdout, the
exit marker and the progress lease never waits on an upload.

The abandon leaves the reference on the row — deliberately, where the
executor's own terminal write clears it — the completion notifier projects
it across the wait boundary, and the node's existing respawn stamps it onto
the fresh task as a transcript ref plus the session id. Only the
conversation survives: the dead host's working tree does not, so the goal
carries the lost-host block, the run records that it resumed from a
checkpoint, and a column names the attempt it took over from.

Checkpointing is an opt-in the agent node sets from the engine's own answer
about its retry policy, so a run whose failure nobody can retry — a
benchmark cell, a review child, a supervisor unit, a single-attempt node —
pays for nothing. Checkpoints are their own retention class with a
two-hour floor: a run writes one a minute, each supersedes the last, and
the growth gate that makes them worth taking is exactly what stops the
content-addressed store deduplicating them.

Supervisor units are out of scope; their respawn can read the same columns
through AgentRetryCauses later.
@ppXD
ppXD force-pushed the feat/resume-a-lost-hosts-run-from-a-durable-checkpoint branch from 9c0d893 to 19a24a0 Compare September 19, 2026 10:41
@ppXD
ppXD merged commit 5121ec6 into main Sep 19, 2026
7 of 9 checks passed
ppXD added a commit that referenced this pull request Sep 23, 2026
#1994 made a lost host's agent continuable, but only for a workflow
agent.run node: the observer checkpoints the live transcript, the
reconciler's abandon keeps the reference, and the node's own respawn
consumes it. A supervisor-spawned unit was excluded — its envelope
never opted in, so its retry restarted the conversation from zero even
though the machinery to continue it was already there.

The staging loop that every spawn wave, retry and resolve passes
through now opts a unit in only when a later retry could consume the
checkpoint: the unit carries a subtask id (resolver units carry none,
and the retry lookup matches on it), and
SupervisorBounds.CanRespawnAfterWave says the run can still afford to
respawn it.

The subtask-attempt lookup falls back to a TERMINAL row's checkpoint
columns when the attempt left no result to read, which is the shape an
abandon leaves. A Running row's checkpoint belongs to a session that
may still be live, so it is never resumed. For "terminal and still
holding a checkpoint" to mean an abandon, a deliberate cancel now
releases the columns as completion does; only the reconciler's abandon
and spool recovery keep them.

The resume fold tells the two kinds of restored conversation apart: a
capture finished with its tree intact, a checkpoint did not, so only
the second carries the lost-host block and its provenance. Its tree
sentence reads the resumed attempt's own pushed branch, not the
effective clone ref, so a dependent unit is never told that its
producer's branch is its own published work. The same change stops a
producer's handoff ref from suppressing the honest-redo line on the
ordinary path when the prior attempt pushed nothing itself.

The billing stays honest: the graded roster keys by subtask and takes
the latest, so one host loss spends one graded attempt, not two.
ppXD added a commit that referenced this pull request Sep 23, 2026
#1994 made a lost host's agent continuable, but only for a workflow
agent.run node: the observer checkpoints the live transcript, the
reconciler's abandon keeps the reference, and the node's own respawn
consumes it. A supervisor-spawned unit was excluded — its envelope
never opted in, so its retry restarted the conversation from zero even
though the machinery to continue it was already there.

The staging loop that every spawn wave, retry and resolve passes
through now opts a unit in only when a later retry could consume the
checkpoint: the unit carries a subtask id (resolver units carry none,
and the retry lookup matches on it), and
SupervisorBounds.CanRespawnAfterWave says the run can still afford to
respawn it.

The subtask-attempt lookup falls back to a TERMINAL row's checkpoint
columns when the attempt left no result to read, which is the shape an
abandon leaves. A Running row's checkpoint belongs to a session that
may still be live, so it is never resumed. For "terminal and still
holding a checkpoint" to mean an abandon, a deliberate cancel now
releases the columns as completion does; only the reconciler's abandon
and spool recovery keep them.

The resume fold tells the two kinds of restored conversation apart: a
capture finished with its tree intact, a checkpoint did not, so only
the second carries the lost-host block and its provenance. Its tree
sentence reads the resumed attempt's own pushed branch, not the
effective clone ref, so a dependent unit is never told that its
producer's branch is its own published work. The same change stops a
producer's handoff ref from suppressing the honest-redo line on the
ordinary path when the prior attempt pushed nothing itself.

The billing stays honest: the graded roster keys by subtask and takes
the latest, so one host loss spends one graded attempt, not two.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant