Repository navigation
feat(harness): keep the transcripts a CLI writes for its subagents - #13
Conversation
What the lab captured of a run was the CLI's stdout, which carries the parent thread. A CLI that delegates writes each subagent's transcript to its own store instead: Muse routinely writes more there than the parent log holds, and Claude puts part of a delegated agent's record on the stream and part in a file, differing from one run to the next. None of it was collected, hashed, or named anywhere, so a run's own account of itself was missing most of the work behind it. Every harness now says where its CLI keeps a session. The run copies that into the artifact directory, hashes each file, and lists them in the manifest, which goes to schema version 2. A copy that could not be taken is recorded as the reason it could not — a CLI that delegated to nobody and a lab that lost what it delegated to are different facts, and only one is worth acting on. Collection never fails a run. The store is read out of the environment the run was spawned with, so the lab and the CLI cannot disagree about where a home is, and a session is found by the id the run was told rather than by re-deriving the CLI's own naming, which would rot silently. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
The Claude CLI files a session under the directory it was run from, so a session resumed elsewhere is written under two projects under the one id. Every located path was copied under its own basename, so the second landed on the first: the manifest listed one transcript, hashed one transcript, and said nothing about the one it had lost. Each copy is now kept where the CLI had it, measured from the CLI's own store, so two projects stay two directories. A path reported from outside that store is kept by name instead of being followed out of the artifact directory.
| } catch (error) { | ||
| return { | ||
| source: store.root, | ||
| files: [], | ||
| gap: SessionTranscriptGaps.FAILED, | ||
| detail: error instanceof Error ? error.message : String(error) | ||
| }; | ||
| } |
There was a problem hiding this comment.
🟡 Medium session-transcript/session-transcript.ts:55
When a later cp or the final hashing throws, collectSessionTranscript returns gap: FAILED with files: [] but leaves the files already copied on disk. With multiple sources, one successful copy followed by a failing source leaves real transcript artifacts in the run directory that are omitted from the manifest, so the manifest no longer describes the directory it attests. Consider removing the directory in the catch block before returning the failure result.
} catch (error) {
+ await rm(directory, { recursive: true, force: true }).catch(() => {});
return {
source: store.root,
files: [],
gap: SessionTranscriptGaps.FAILED,
detail: error instanceof Error ? error.message : String(error)
};🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @packages/harness/src/session-transcript/session-transcript.ts around lines 55-62:
When a later `cp` or the final hashing throws, `collectSessionTranscript` returns `gap: FAILED` with `files: []` but leaves the files already copied on disk. With multiple sources, one successful copy followed by a failing source leaves real transcript artifacts in the run directory that are omitted from the manifest, so the manifest no longer describes the directory it attests. Consider removing the `directory` in the `catch` block before returning the failure result.
…ipts # Conflicts: # packages/harness/src/claude-cli/claude-harness.test.ts # packages/harness/src/claude-cli/claude-harness.ts # packages/harness/src/cli-agent-harness/cli-agent-harness.ts # packages/harness/src/cli-agent-harness/harness-run-artifacts.ts # packages/harness/src/codex-cli/codex-harness.ts # packages/harness/src/deepseek-cli/deepseek-harness.ts # packages/harness/src/glm-cli/glm-harness.ts # packages/harness/src/muse-cli/muse-harness.ts
| @@ -347,6 +357,11 @@ export abstract class CliAgentHarness implements AgentHarness { | |||
|
|
|||
| await writeStderrArtifact(files.stderrPath, processExit.stderr); | |||
There was a problem hiding this comment.
🟠 High cli-agent-harness/cli-agent-harness.ts:358
If the CLI process exits within timeoutMs but collectSessionTranscript takes long enough to push the total elapsed time past the deadline, watchdog.timedOut() returns true and determineRunStatus records the run as timed out even though the process completed successfully. The watchdog signal is still live during transcript collection because it is only consumed when processCompleted is awaited, so any collection overrun flips the status to a timeout that never happened. Stop the watchdog before calling collectSessionTranscript, or exclude post-run collection from the timeout status check.
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @packages/harness/src/cli-agent-harness/cli-agent-harness.ts around line 358:
If the CLI process exits within `timeoutMs` but `collectSessionTranscript` takes long enough to push the total elapsed time past the deadline, `watchdog.timedOut()` returns true and `determineRunStatus` records the run as timed out even though the process completed successfully. The watchdog signal is still live during transcript collection because it is only consumed when `processCompleted` is awaited, so any collection overrun flips the status to a timeout that never happened. Stop the watchdog before calling `collectSessionTranscript`, or exclude post-run collection from the timeout status check.
What Changed
Every harness now says where its CLI keeps its own record of a session. At the end of a run the lab copies that into the artifact directory, hashes each file, and lists them in the manifest under
artifacts.session. The manifest schema goes to version 2.sessionStore()is abstract, not defaulted. A sixth harness cannot be added without answering the question, because a harness that stayed silent would lose transcripts with nothing saying so.When there is no copy, the manifest says which of four things happened rather than showing an empty list:
no_session,no_store,not_found,failed. Collection never fails a run — a run that produced an answer produced it whether or not the lab could read the CLI's account of it afterwards.Why
What the lab captured was the CLI's stdout, which carries the parent thread. A CLI that delegates writes each subagent's transcript to its own store and nowhere else. That is most of the work behind a research run, and none of it was collected, hashed, or named anywhere.
Measured against the live CLIs on this machine, not from documentation:
Muse Code 0.1.0 — one real session from this machine's store: 74 subagents, 8.8 MB of transcripts against a 6.0 MB parent log. More than half the session was outside anything the lab held. The repository already knew:
muse-run-arguments.ts:14-17explains that the session log is deliberately left switched on because "a Muse run delegates to subagents whose transcripts never appear there — they are written beside the session as files of their own, and turning the log off would throw away the only copy." It was left at that.Claude Code 2.1.224 — two probe runs, one foreground subagent and one background, each with a transcript of 11 entries:
Not only partial — partial differently each time. No
stream_eventcarried aparent_tool_use_idin either run, so a delegated agent's reasoning deltas never reach the stream at all.Codex 0.146 — the locator is written and tested, but the ChatGPT quota on this machine is spent until Aug 9, so its subagent file layout is not confirmed by a live delegating run. What is confirmed:
multi_agentis a stable feature on by default,--ignore-user-configdoes not disable it, and a rollout is written for every thread including one that died on quota. Flagging this rather than implying coverage I do not have.Two decisions worth arguing with
Copy, don't hash in place. A hash of a file in
~/.claudeattests to something outside the lab, pruned on the vendor's schedule, that a purge cannot reach. The README promises a lab is a directory you copy to keep and delete to be rid of; that only holds if the bytes are in it. The cost is disk — a Muse run can add ~9 MB — which is the same order as thenative-events.jsonlalready stored beside it. No cap; say so if you want one.Find the session by id, not by re-deriving the CLI's own naming. Claude's project directory is a slug of the run's cwd and the derivation is entirely the CLI's to change. A lab built on it would stop collecting silently the day it did. The session id is something the run was told.
The store is read out of the environment the run was actually spawned with, through each CLI's own variable —
CLAUDE_CONFIG_DIR,CODEX_HOME,XDG_DATA_HOME, all three verified present in the shipped binaries — so the lab and the CLI cannot disagree about where a home is.Verification
Beyond the unit tests, a live end-to-end run through
ClaudeHarnessagainst the realclaudebinary, with a real subagent:Those 12,755 bytes are what used to be lost. The copy's digest was checked independently with
shasum -a 256against the original still sitting in~/.claude— identical.pnpm check(lint, typecheck, 9 workspaces of tests, build) passes.Merge #12 first if you can — it fixes a pre-existing flake that makes this branch's CI fail at random. It is independent of this work and does not block it.
Checklist
Claude Opus 5 (1M context) via Claude Code.
🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmithwith what you need. Autofix is disabled.Note
Collect and attach CLI session transcripts (including subagent transcripts) to harness run artifacts
collectSessionTranscriptutility that locates a CLI's session files after a run, copies them into the artifact directory, hashes and normalizes permissions, and attaches them toartifacts.session; failures are recorded as gaps without failing the run.claudeSessionStore), Codex (codexSessionStore), and Muse (museSessionStore), each resolving the store root from environment variables or HOME fallback paths.sessionStore(environment)method toCliAgentHarnessthat all concrete harnesses (ClaudeHarness,CodexHarness,DeepseekHarness,GlmHarness,MuseHarness) now implement.HARNESS_MANIFEST_SCHEMA_VERSIONfrom 1 to 2 and addssession: SessionTranscriptas a required field onHarnessArtifactsandNonManifestArtifacts.HarnessArtifactsorNonManifestArtifactsmust now supply asessionfield; manifests will declare schema version 2.Macroscope summarized 69c2c76.