Skip to content

feat(harness): keep the transcripts a CLI writes for its subagents - #13

Merged
dibenkobit merged 3 commits into
mainfrom
feat/subagent-transcripts
Aug 7, 2026
Merged

dibenkobit merged 3 commits into
mainfrom
feat/subagent-transcripts

Conversation

@dibenkobit

@dibenkobit dibenkobit commented Aug 7, 2026 •

Copy link
Copy Markdown
Member

What Changed

Every harness now says where its CLI keeps its own record of a session. At the end of a run the lab copies that into the artifact directory, hashes each file, and lists them in the manifest under artifacts.session. The manifest schema goes to version 2.

packages/harness/src/session-transcript/     the collector, its contracts and its gaps
packages/harness/src/claude-cli/claude-session-store.ts
packages/harness/src/codex-cli/codex-session-store.ts
packages/harness/src/muse-cli/muse-session-store.ts

sessionStore() is abstract, not defaulted. A sixth harness cannot be added without answering the question, because a harness that stayed silent would lose transcripts with nothing saying so.

When there is no copy, the manifest says which of four things happened rather than showing an empty list: no_session, no_store, not_found, failed. Collection never fails a run — a run that produced an answer produced it whether or not the lab could read the CLI's account of it afterwards.

Why

What the lab captured was the CLI's stdout, which carries the parent thread. A CLI that delegates writes each subagent's transcript to its own store and nowhere else. That is most of the work behind a research run, and none of it was collected, hashed, or named anywhere.

Measured against the live CLIs on this machine, not from documentation:

Muse Code 0.1.0 — one real session from this machine's store: 74 subagents, 8.8 MB of transcripts against a 6.0 MB parent log. More than half the session was outside anything the lab held. The repository already knew: muse-run-arguments.ts:14-17 explains that the session log is deliberately left switched on because "a Muse run delegates to subagents whose transcripts never appear there — they are written beside the session as files of their own, and turning the log off would throw away the only copy." It was left at that.

Claude Code 2.1.224 — two probe runs, one foreground subagent and one background, each with a transcript of 11 entries:

entries on stdout lost
foreground subagent 7 of 11 thinking block, final text, 2 attachments
background subagent 8 of 11 the subagent's own prompt, 2 attachments

Not only partial — partial differently each time. No stream_event carried a parent_tool_use_id in either run, so a delegated agent's reasoning deltas never reach the stream at all.

Codex 0.146 — the locator is written and tested, but the ChatGPT quota on this machine is spent until Aug 9, so its subagent file layout is not confirmed by a live delegating run. What is confirmed: multi_agent is a stable feature on by default, --ignore-user-config does not disable it, and a rollout is written for every thread including one that died on quota. Flagging this rather than implying coverage I do not have.

Two decisions worth arguing with

Copy, don't hash in place. A hash of a file in ~/.claude attests to something outside the lab, pruned on the vendor's schedule, that a purge cannot reach. The README promises a lab is a directory you copy to keep and delete to be rid of; that only holds if the bytes are in it. The cost is disk — a Muse run can add ~9 MB — which is the same order as the native-events.jsonl already stored beside it. No cap; say so if you want one.

Find the session by id, not by re-deriving the CLI's own naming. Claude's project directory is a slug of the run's cwd and the derivation is entirely the CLI's to change. A lab built on it would stop collecting silently the day it did. The session id is something the run was told.

The store is read out of the environment the run was actually spawned with, through each CLI's own variable — CLAUDE_CONFIG_DIR, CODEX_HOME, XDG_DATA_HOME, all three verified present in the shipped binaries — so the lab and the CLI cannot disagree about where a home is.

Verification

Beyond the unit tests, a live end-to-end run through ClaudeHarness against the real claude binary, with a real subagent:

status          : succeeded
schemaVersion   : 2
session.source  : /Users/dibenkobit/.claude
session.gap     : (none)
collected files :
    17654  df510b918300  …/session/f09f2a26….jsonl
    12755  df5e32f72804  …/session/f09f2a26…/subagents/agent-a79e6381755ff4322.jsonl
      140  f63089bde567  …/session/f09f2a26…/subagents/agent-a79e6381755ff4322.meta.json

Those 12,755 bytes are what used to be lost. The copy's digest was checked independently with shasum -a 256 against the original still sitting in ~/.claude — identical.

pnpm check (lint, typecheck, 9 workspaces of tests, build) passes.

Merge #12 first if you can — it fixes a pre-existing flake that makes this branch's CI fail at random. It is independent of this work and does not block it.

Checklist

  • This PR is small and focused
  • I explained what changed and why
  • I included screenshots for any UI changes, with a before only where one existed — no UI changes; nothing here is rendered yet
  • I included a video for animation/interaction changes — not applicable

Claude Opus 5 (1M context) via Claude Code.

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith with what you need. Autofix is disabled.

Note

Collect and attach CLI session transcripts (including subagent transcripts) to harness run artifacts

  • Adds a collectSessionTranscript utility that locates a CLI's session files after a run, copies them into the artifact directory, hashes and normalizes permissions, and attaches them to artifacts.session; failures are recorded as gaps without failing the run.
  • Introduces session store implementations for Claude (claudeSessionStore), Codex (codexSessionStore), and Muse (museSessionStore), each resolving the store root from environment variables or HOME fallback paths.
  • Adds an abstract sessionStore(environment) method to CliAgentHarness that all concrete harnesses (ClaudeHarness, CodexHarness, DeepseekHarness, GlmHarness, MuseHarness) now implement.
  • Bumps HARNESS_MANIFEST_SCHEMA_VERSION from 1 to 2 and adds session: SessionTranscript as a required field on HarnessArtifacts and NonManifestArtifacts.
  • Behavioral Change: any code constructing HarnessArtifacts or NonManifestArtifacts must now supply a session field; manifests will declare schema version 2.

Macroscope summarized 69c2c76.

What the lab captured of a run was the CLI's stdout, which carries the
parent thread. A CLI that delegates writes each subagent's transcript to
its own store instead: Muse routinely writes more there than the parent
log holds, and Claude puts part of a delegated agent's record on the
stream and part in a file, differing from one run to the next. None of it
was collected, hashed, or named anywhere, so a run's own account of
itself was missing most of the work behind it.

Every harness now says where its CLI keeps a session. The run copies that
into the artifact directory, hashes each file, and lists them in the
manifest, which goes to schema version 2. A copy that could not be taken
is recorded as the reason it could not — a CLI that delegated to nobody
and a lab that lost what it delegated to are different facts, and only
one is worth acting on. Collection never fails a run.

The store is read out of the environment the run was spawned with, so the
lab and the CLI cannot disagree about where a home is, and a session is
found by the id the run was told rather than by re-deriving the CLI's own
naming, which would rot silently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
openlab-landing Ready Ready Preview Aug 7, 2026 4:11pm

Request Review

Comment thread packages/harness/src/session-transcript/session-transcript.ts
The Claude CLI files a session under the directory it was run from, so a session
resumed elsewhere is written under two projects under the one id. Every located
path was copied under its own basename, so the second landed on the first: the
manifest listed one transcript, hashed one transcript, and said nothing about the
one it had lost.

Each copy is now kept where the CLI had it, measured from the CLI's own store, so
two projects stay two directories. A path reported from outside that store is
kept by name instead of being followed out of the artifact directory.
Comment on lines +55 to +62
} catch (error) {
return {
source: store.root,
files: [],
gap: SessionTranscriptGaps.FAILED,
detail: error instanceof Error ? error.message : String(error)
};
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium session-transcript/session-transcript.ts:55

When a later cp or the final hashing throws, collectSessionTranscript returns gap: FAILED with files: [] but leaves the files already copied on disk. With multiple sources, one successful copy followed by a failing source leaves real transcript artifacts in the run directory that are omitted from the manifest, so the manifest no longer describes the directory it attests. Consider removing the directory in the catch block before returning the failure result.

    } catch (error) {
+        await rm(directory, { recursive: true, force: true }).catch(() => {});
        return {
            source: store.root,
            files: [],
            gap: SessionTranscriptGaps.FAILED,
            detail: error instanceof Error ? error.message : String(error)
        };
🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @packages/harness/src/session-transcript/session-transcript.ts around lines 55-62:

When a later `cp` or the final hashing throws, `collectSessionTranscript` returns `gap: FAILED` with `files: []` but leaves the files already copied on disk. With multiple sources, one successful copy followed by a failing source leaves real transcript artifacts in the run directory that are omitted from the manifest, so the manifest no longer describes the directory it attests. Consider removing the `directory` in the `catch` block before returning the failure result.

…ipts

# Conflicts:
#	packages/harness/src/claude-cli/claude-harness.test.ts
#	packages/harness/src/claude-cli/claude-harness.ts
#	packages/harness/src/cli-agent-harness/cli-agent-harness.ts
#	packages/harness/src/cli-agent-harness/harness-run-artifacts.ts
#	packages/harness/src/codex-cli/codex-harness.ts
#	packages/harness/src/deepseek-cli/deepseek-harness.ts
#	packages/harness/src/glm-cli/glm-harness.ts
#	packages/harness/src/muse-cli/muse-harness.ts
@dibenkobit
dibenkobit merged commit f6be6d1 into main Aug 7, 2026
7 checks passed
@@ -347,6 +357,11 @@ export abstract class CliAgentHarness implements AgentHarness {

await writeStderrArtifact(files.stderrPath, processExit.stderr);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High cli-agent-harness/cli-agent-harness.ts:358

If the CLI process exits within timeoutMs but collectSessionTranscript takes long enough to push the total elapsed time past the deadline, watchdog.timedOut() returns true and determineRunStatus records the run as timed out even though the process completed successfully. The watchdog signal is still live during transcript collection because it is only consumed when processCompleted is awaited, so any collection overrun flips the status to a timeout that never happened. Stop the watchdog before calling collectSessionTranscript, or exclude post-run collection from the timeout status check.

🚀 Reply "fix it for me" or copy this AI Prompt for your agent:
In file @packages/harness/src/cli-agent-harness/cli-agent-harness.ts around line 358:

If the CLI process exits within `timeoutMs` but `collectSessionTranscript` takes long enough to push the total elapsed time past the deadline, `watchdog.timedOut()` returns true and `determineRunStatus` records the run as timed out even though the process completed successfully. The watchdog signal is still live during transcript collection because it is only consumed when `processCompleted` is awaited, so any collection overrun flips the status to a timeout that never happened. Stop the watchdog before calling `collectSessionTranscript`, or exclude post-run collection from the timeout status check.

This branch was successfully deployed

1 active deployment
Preview — 69c2c76a Deployed Aug 7, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant