Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
-- 0234_agent_run_session_transcript_checkpoint.sql
--
-- A run whose launch host dies is NOT continued by another host, because nothing it produced mid-run is durable.
--
-- The resumable CLI session transcript — the file the harness's own `--resume` reads — is captured exactly once, at
-- run END, from the LAUNCHING host's spool (AgentRunExecutor.CaptureSessionTranscriptAsync reads
-- LocalProcessRunner.ConfigHomePath(handle.SpoolDirectory)). When the host holding that spool is gone, the file is
-- gone with it: AgentRunReconcilerService abandons the run (a foreign handle is left alone until its own deadline,
-- then failed) and every reader is left with a cold restart as the only recovery.
--
-- The retry itself already happens: the reconciler abandons the run, the agent.run node grades that as a retryable
-- failure, and the node's retry policy stages a fresh agent. It is the CONVERSATION that is lost, so the retry is
-- cold. These columns are what make it warm. The observer's own drain tick uploads the live transcript to the
-- artifact store and stamps the reference here; the reconciler's abandon leaves it on the row (where the executor's
-- own terminal write clears it), the completion notifier projects it across the wait boundary, and the node's
-- respawn stamps it onto the fresh attempt's task.
--
-- Shape, and why each part of it:
--
-- * NULLABLE with NO DEFAULT. agent_run is a hot table; a nullable ADD COLUMN with no default is catalog-only in
-- PostgreSQL (no table rewrite, no long ACCESS EXCLUSIVE hold), which a DEFAULT would not be on every supported
-- server. NULL is also the escape for the rollback direction: an older binary that has never heard of these
-- columns simply does not select them, and a row they were never written to reads as "no checkpoint" — which is
-- precisely the pre-3c behaviour, a cold start.
--
-- * The artifact id is a SOFT link, like every other *_artifact_id in this schema — no FK to workflow_artifact.
-- It is registered in ArtifactReferenceOracle.ReferenceSites (test-pinned, and cross-checked against the EF model
-- by ArtifactReferenceOracleTests) so the reaper can never collect a checkpoint a live run still names.
--
-- * A PARTIAL INDEX on that soft link, for the reason 0144 gives for the ten it added with the oracle: the oracle
-- asks "does ANY row anywhere still reference this artifact id" as one EXISTS per site, and an unindexed site on
-- a run-scale table makes that EXISTS a sequential scan of agent_run on every artifact the reaper claims. Partial
-- on IS NOT NULL because a checkpoint is the exception, so the index stays small on the hot table. It also builds
-- instantly here: the column is NULL on every existing row, so there is nothing to insert into it.
--
-- LOCKS, stated as 0144 states them: DbUp runs the whole upgrade in ONE transaction, so the CREATE INDEX below holds
-- its SHARE lock on agent_run (blocking writes, not reads) until that transaction commits. CREATE INDEX CONCURRENTLY
-- cannot run inside a transaction, so shortening that window is a change to DbUpRunner, not to this file.
--
-- * A THIRD column, resumed_from_agent_run_id, so a continuation NAMES the run it took over from in a column
-- rather than only in the sentence an event carries. Same shape as every other agent_run cross-reference: a soft
-- link with no FK (agent runs are managed independently of one another), nullable, no default, no index — nothing
-- queries by it yet, and adding an index for a reader that does not exist would cost every agent_run write. It is
-- NOT an artifact id, so the reference oracle correctly does not probe it.
--
-- Rollback: ALTER TABLE agent_run DROP COLUMN resumed_from_agent_run_id, DROP COLUMN
-- session_transcript_checkpoint_at, DROP COLUMN session_transcript_checkpoint_artifact_id (the index goes with the
-- column); the stored transcripts then age out under their own retention declaration
-- (ArtifactRetentionClass.SessionTranscriptCheckpoint) once nothing references them.

ALTER TABLE agent_run
ADD COLUMN session_transcript_checkpoint_artifact_id uuid NULL,
ADD COLUMN session_transcript_checkpoint_at timestamptz NULL,
ADD COLUMN resumed_from_agent_run_id uuid NULL;

CREATE INDEX ix_agent_run_session_transcript_checkpoint_artifact
ON agent_run (session_transcript_checkpoint_artifact_id)
WHERE session_transcript_checkpoint_artifact_id IS NOT NULL;

COMMENT ON COLUMN agent_run.session_transcript_checkpoint_artifact_id IS
'Soft link to workflow_artifact.id holding this run''s resumable session transcript as of its most recent mid-run checkpoint — what another host continues from when this run''s host dies. NULL until the first checkpoint.';

COMMENT ON COLUMN agent_run.session_transcript_checkpoint_at IS
'When session_transcript_checkpoint_artifact_id was taken: the cadence clock for the next checkpoint, and the age that says how much of the conversation survived the lost host.';

COMMENT ON COLUMN agent_run.resumed_from_agent_run_id IS
'The abandoned agent run this one succeeded after a host loss, and was staged to continue from its session checkpoint. Whether that checkpoint could be read is only known at launch, so this records succession, not recovery. Soft link (no FK, like every other agent_run cross-reference); NULL for every ordinary run.';
33 changes: 31 additions & 2 deletions backend/src/CodeSpace.Core/Persistence/Entities/AgentRun.cs
Original file line number Diff line number Diff line change
Expand Up @@ -67,11 +67,40 @@ public class AgentRun : IEntity<Guid>, IAuditable
/// <summary>
/// P3.1a: the harness-native session/thread id captured off the run's CLI conversation (Claude's
/// <c>session_id</c>, Codex's <c>thread_id</c>) — promoted from <c>result_jsonb</c> to a first-class column so a
/// rerun's CONTINUE lookup is a column read, not a JSON probe. NULL while in-flight, and for a run whose stream
/// carried no session id (a pre-session CLI). Set on completion from <c>AgentRunResult.SessionId</c>.
/// rerun's CONTINUE lookup is a column read, not a JSON probe. NULL for a run whose stream carried no session id
/// (a pre-session CLI). Set on completion from <c>AgentRunResult.SessionId</c>, and — since 3c — already at the
/// run's first session-transcript checkpoint, because a checkpoint nothing can ADDRESS is not resumable: the
/// continuation needs this id to hand the CLI its <c>--resume</c>. Every reader that treats a non-null id as
/// "resumable" also requires a transcript out of <see cref="ResultJson"/> (<c>TryResumable</c>, both-or-neither),
/// so an in-flight row carrying one is skipped exactly as a null one was.
/// </summary>
public string? SessionId { get; set; }

/// <summary>
/// 3c: the artifact holding this run's resumable session transcript as of its most recent MID-RUN checkpoint —
/// the only thing a run leaves behind that another host can continue from after the launching host dies. NULL
/// until the first checkpoint, and for every run whose harness has no addressable session transcript. A soft link
/// to <c>workflow_artifact.id</c>, and it is exactly that: the reference the artifact reaper's oracle probes
/// (<c>ArtifactReferenceOracle.ReferenceSites</c>) so a live checkpoint is never collected.
/// </summary>
public Guid? SessionTranscriptCheckpointArtifactId { get; set; }

/// <summary>When <see cref="SessionTranscriptCheckpointArtifactId"/> was taken. The cadence clock for the next checkpoint, and the age an operator (or a resumed attempt) reads to know how much of the conversation survived the lost host.</summary>
public DateTimeOffset? SessionTranscriptCheckpointAt { get; set; }

/// <summary>
/// 3c: the abandoned run this one was STAGED to continue from a session checkpoint — the structured link, so
/// "which attempt took over from which" is a column rather than a sentence buried in an event. Soft link (no
/// FK), like every other agent-run cross-reference. NULL for every ordinary run.
///
/// <para>Staged, not necessarily restored: this is written when the row is created, and whether the checkpoint
/// could actually be READ is only knowable at launch (a reaped blob, an unreachable destination — then the
/// attempt runs cold and says so). The honest reading of a non-null value is therefore "this attempt succeeded
/// that one after it lost its host", which is true either way; what was recovered is the confinement record's
/// <c>ResumedFromCheckpointAt</c>, which the degrade clears.</para>
/// </summary>
public Guid? ResumedFromAgentRunId { get; set; }

/// <summary>Worker liveness ping; a stuck-Running reconciler reads this to recover crashed runs.</summary>
public DateTimeOffset? HeartbeatAt { get; set; }

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,13 @@ public void Configure(EntityTypeBuilder<AgentRun> builder)
// is a column read. Nullable (a pre-session CLI / in-flight run has none).
builder.Property(r => r.SessionId).HasColumnName("session_id");

// 3c (migration 0234) — the mid-run session-transcript checkpoint. Nullable with no default: the column is
// metadata-only on this hot table, and NULL is the escape an older binary and every pre-3c run reads as
// "no checkpoint, cold start".
builder.Property(r => r.SessionTranscriptCheckpointArtifactId).HasColumnName("session_transcript_checkpoint_artifact_id");
builder.Property(r => r.SessionTranscriptCheckpointAt).HasColumnName("session_transcript_checkpoint_at");
builder.Property(r => r.ResumedFromAgentRunId).HasColumnName("resumed_from_agent_run_id");

builder.Property(r => r.RunnerHandleJson).HasColumnName("runner_handle").HasColumnType("jsonb");
builder.Property(r => r.SpoolCleanupAttempts).HasColumnName("spool_cleanup_attempts");
builder.Property(r => r.SpoolCleanupLastAttemptAt).HasColumnName("spool_cleanup_last_attempt_at");
Expand Down
46 changes: 46 additions & 0 deletions backend/src/CodeSpace.Core/Services/Agents/AgentRetryContinuity.cs
Original file line number Diff line number Diff line change
Expand Up @@ -15,4 +15,50 @@ public static class AgentRetryContinuity

/// <summary>Append <see cref="HonestNoContinuityHint"/> to a resumed task's goal. One composition, so the two lanes cannot drift on the separator either.</summary>
public static string WithHonestNoContinuityHint(string goal) => $"{goal}\n\n{HonestNoContinuityHint}";

/// <summary>
/// 3c: what a CROSS-HOST continuation is told, and the reason it needs its own sentence. The two lanes above
/// retry an attempt that FINISHED on a live host, so their only open question is whether a branch was pushed.
/// This lane continues an attempt whose machine was lost mid-run: the conversation comes from a checkpoint taken
/// some time before the loss, and the working tree is simply gone. Both halves have to be said, because the
/// restored transcript will describe edits — possibly edits made after the checkpoint — that the new workspace
/// does not contain, and an agent that is not told will read its own transcript as evidence about files it cannot
/// see.
/// </summary>
public const string LostHostPreamble = "Note: the machine running your previous attempt was lost mid-run. Your conversation is restored from a checkpoint taken before that, so it may describe work you did after the checkpoint, and it may be missing your last few turns.";

/// <summary>Said when the lost attempt HAD published a branch: the workspace is checked out at it, so the published work is present and only the unpublished remainder is gone. Takes the branch name so the agent can verify rather than take the claim on trust.</summary>
public static string LostHostPublishedBranchHint(string branch) =>
$"Your previous attempt published branch `{branch}`, and this workspace is checked out AT that branch — that work is here. Anything you had NOT published to it died with the machine, so check the files before continuing and redo whatever is missing.";

/// <summary>
/// Append the cross-host continuation's honesty block to a resumed task's goal: the preamble always, then what
/// is true about the TREE.
///
/// <para>Three cases, and the third is why this takes a flag rather than inferring from the branch alone. A
/// published branch means the pushed work IS here and only the unpublished remainder died. No branch on a
/// repository-backed run means nothing was preserved, which is exactly <see cref="HonestNoContinuityHint"/> —
/// the sentence the other two lanes already say, reused verbatim so "your prior work is not here" can never mean
/// two different things. And a run with NO repository (<paramref name="treeOwed"/> false) had no working tree to
/// lose, so it gets the preamble alone: telling an analysis-only agent to redo file changes would be a claim
/// about a git fact its run never had.</para>
/// </summary>
public static string WithLostHostHint(string goal, string? publishedBranch, bool treeOwed) =>
$"{goal}\n\n{LostHostPreamble}{TreeStateSentence(publishedBranch, treeOwed)}";

/// <summary>
/// 3c: said when the lost host's checkpoint could not be READ — reaped, or its storage unreachable. The attempt
/// still runs (failing it would spend the retry this whole path exists to improve), but it runs COLD, and an
/// agent that was going to be handed a conversation must be told it is not getting one. Appended to the
/// lost-host block rather than replacing it: the machine really was lost, which is still the reason.
/// </summary>
public const string LostHostCheckpointUnreadableHint = "Your previous conversation could not be recovered either — the checkpoint it was stored in is no longer readable — so you are starting this task from the beginning.";

/// <summary>Replace a resumed goal's promise of a restored conversation with the truth that there is none. Used when the checkpoint ref resolves to nothing at launch, which is the only moment that fact is knowable.</summary>
public static string WithUnreadableCheckpointHint(string goal) => $"{goal}\n\n{LostHostCheckpointUnreadableHint}";

private static string TreeStateSentence(string? publishedBranch, bool treeOwed) =>
!treeOwed ? ""
: string.IsNullOrWhiteSpace(publishedBranch) ? $" {HonestNoContinuityHint}"
: $" {LostHostPublishedBranchHint(publishedBranch)}";
}
Loading
Loading