Skip to content

feat(hooks): export a post-admission execution receipt to tool_call_after (#6689) - #6713

Merged
Hmbown merged 7 commits into
mainfrom
feat/6689-hook-execution-receipt
Sep 29, 2026
Merged

Hmbown merged 7 commits into
mainfrom
feat/6689-hook-execution-receipt

Conversation

@Hmbown

@Hmbown Hmbown commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Closes #6689. Refs #6582; that issue remains open because this uses the proposed environment variable rather than a stdin JSON envelope, with state and output_kind field names.

After an admitted foreground shell command runs, tool_call_after and failed-call on_error hooks receive a schema-1 JSON DEEPSEEK_TOOL_EXECUTION_RECEIPT. It carries the command from the actual spawn record, the directory passed to the process, completion state, observed exit code, and bounded stdout/stderr previews. Both the TUI and Runtime API use the existing HookContext::with_tool_outcome boundary; the Runtime API strips this hook-only metadata before persisting or emitting the tool item.

The directory is resolved asynchronously before spawning and that same path supplies process and receipt identity. A command that retargets a workspace symlink cannot change the reported starting directory. Both foreground and background hook launchers clear an inherited receipt before overlaying the current call's context, so an absent or oversized current receipt cannot expose another call's receipt.

Receipts remain exact or absent: oversized or invalid identities, unresolved directories, failed waits, PTY/interactive, background/detached, sandboxed/external-backend, read-only argv, PowerShell and Windows paths do not claim an exact receipt. Preview limits apply to the tool's retained output; earlier process output may already have been dropped by its bounded capture. A too-large or unsupported-schema document is omitted intact rather than truncated. Hooks gain observation only, with no new admission or execution authority.

The contract and test design follow @wuisabel-gif's reference branch, wuisabel-gif/CodeWhale@feat/tool-call-after-execution-receipt. The original implementation adapts that design to the current shared hook boundary and process-manager record, includes failed lowercase bash calls, and preserves Isabel Wu's co-author attribution and release credits. Documentation is updated in English and Simplified Chinese.

Validation:

  • Focused codewhale-tui tests at repair checkpoint 69eebf8ae: 7 passed, 0 failed. These execute a symlink-retargeting shell command and six real hook children, covering absent, oversized and current receipts in both launch modes.
  • Latest ordinary merge of main 47a014c6 is da1212d39ba365407970c0579e14b7d036c43a05. npm test: 670 passed, 0 failed (68 CLI, 16 runtime SDK, 50 extension host, 536 web). npm run check:web: passed, including production web build.
  • Version consistency passed. After the merge commit, all eight contributors in the release range have matching credits on all three surfaces.
  • Root and independent reviews verified that all seven original hook implementation, test and documentation patches are preserved exactly after normalization. Only four changelog/credit conflicts required manual resolution; the Runtime metadata exclusion remains intact. Formatting and diff checks passed. No new local Rust build was needed for those documentation conflicts.

The first child npm attempt could not bind its local HTTP fixtures; the root retry passed. Hosted CI on the latest head remains required. Local fixtures do not establish provider, native, deployment or release qualification.

…fter (#6689)

tool_call_after hooks could not tell what a shell tool actually ran. They
got the result text, the success flag, the exit code and the status, but not
the command or its directory. A tool_call_before rewrite can change the
before-hook input they would otherwise re-derive those from.

DEEPSEEK_TOOL_EXECUTION_RECEIPT now carries schema-1 JSON with these fields:
command, cwd, state (completed/interrupted), scope, exit_code (observed or
null), bounded stdout/stderr previews, truncation flags and output_kind.

- Identity reuses the spawn record. execute_foreground_via_background reads
  command and working_dir from the BackgroundShell record the process manager
  created, under the same lock that marks the run as foreground.
- HookContext::with_tool_outcome reads metadata.execution_receipt from Ok
  and Err results. That one seam covers the TUI and Runtime API completion
  hooks, and on_error for a failed shell call. Lowercase bash failures carry
  the receipt too.
- Exact or absent. An identity over 8 KiB, containing NUL, or with a
  relative or non-UTF-8 directory gets no receipt. Previews keep both ends
  and halve until the serialized JSON is at most 32 KiB. The hook boundary
  drops a non-schema-1 or oversized receipt instead of truncating it.
- Scope is settled, pipe-backed, unsandboxed, local foreground runs only. No
  receipt for background, moved-to-/jobs, PTY/interactive, OS sandbox,
  external backend, read-only argv hardening, Windows, or pre-exec refusals.
- Documented in docs/HOOKS.md and docs/zh_hans/HOOKS.md.

The contract and test design follow Isabel Wu's reference branch
wuisabel-gif/CodeWhale@feat/tool-call-after-execution-receipt. This version
is re-implemented on the current with_tool_outcome seam and on main's
process-manager record.

Evidence (macOS, CARGO_TARGET_DIR outside the repo):
  cargo test -p codewhale-tui --lib -- execution_receipt
    -> test result: ok. 3 passed; 0 failed; 0 ignored; 13816 filtered out
  same run with the metadata read and identity capture disabled
    -> test result: FAILED. 1 passed; 2 failed (both new behavior tests fail)
  cargo test -p codewhale-tui --lib -- hooks:: tools::shell:: tool_routing
    -> test result: ok. 340 passed; 0 failed; 1 ignored
  cargo test -p codewhale-tui --lib -- runtime_tool_completion_fires_after_and_error_hooks runtime_shell_completion_delivers_exit_code_and_status_to_hooks
    -> test result: ok. 2 passed; 0 failed
  cargo clippy -p codewhale-tui --lib --tests --locked -- -D warnings (CI allow-list)
    -> Finished, 0 warnings
  cargo fmt --all -- --check -> clean

Co-authored-by: Isabel Wu <231155141+wuisabel-gif@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Copilot AI balanced review requested due to automatic review settings September 29, 2026 03:22

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

Devin Review

Comment on lines +5025 to +5026
let (stdout, stdout_clipped) = execution_receipt_preview(&result.stdout, budget);
let (stderr, stderr_clipped) = execution_receipt_preview(&result.stderr, budget);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Receipt previews inherit earlier output clipping

For long runs, shell_execution_receipt previews the already-bounded ShellResult, not the raw streams. The lowercase bash accumulator can retain only a tail, so its receipt cannot show the original first bytes. The documented preview contract may need to distinguish tool output from raw process output.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Hmbown and others added 3 commits September 28, 2026 21:26
Review follow-ups on the tool_call_after execution receipt:

- cwd has one spelling. An explicit cwd reached the spawn already
  canonicalized by ToolContext::resolve_path, while the default workspace
  was passed as the session opened it, so a workspace opened through a
  symlink reported two paths for one directory. The receipt now
  canonicalizes the recorded directory with tokio::fs::canonicalize (off
  the runtime thread); a directory that no longer resolves gets no receipt.
  The test opens the workspace through a symlink and compares exact strings
  instead of re-canonicalizing them.
- A wait error is not an interruption. BackgroundShell::poll marks a failed
  try_wait (wait_failed); such a run gets no receipt instead of state
  "interrupted" for a process that may have exited normally.
- PowerShell is excluded on every platform. $SHELL=pwsh on Unix makes the
  dispatcher wrap the source or run it from a temp -File, so the receipt's
  command would not be what ran.
- Hook-only. The receipt is built only while a tool_call_after or on_error
  hook is registered, and the Runtime API strips it before persisting and
  emitting the item, which already carries the output.
- docs/HOOKS.md and docs/zh_hans/HOOKS.md describe all four.

Evidence (macOS, CARGO_TARGET_DIR outside the repo):
  cargo test -p codewhale-tui --lib --locked -- execution_receipt
    -> test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 13816 filtered out
  same test with the canonicalize step removed
    -> test result: FAILED. 0 passed; 1 failed (left: .../link, right: /private/var/.../real)
  cargo test -p codewhale-tui --lib --locked -- hooks:: tools::shell:: tool_routing runtime_tool_completion runtime_shell_completion
    -> test result: ok. 341 passed; 0 failed; 1 ignored; 0 measured; 13477 filtered out
  cargo clippy -p codewhale-tui --lib --tests --locked -- -D warnings (CI allow-list)
    -> Finished, 0 warnings
  cargo fmt --all -- --check -> clean

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…6689)

b965cd7 carries a Co-authored-by trailer for Isabel Wu, whose reference
branch set the tool_call_after execution-receipt contract and test design.
check-contributor-credit.py (Version drift job) failed because the handle
was not in web/lib/release-credits.ts. Add her to the v0.10.1 band on every
credit surface so the parity tests stay exact: RELEASE_CONTRIBUTORS, the
0.10.1 CHANGELOG Contributors block (and the synced crates/tui copy),
the docs/CONTRIBUTORS.md v0.10.1 band, and requiredCandidateCredits.

Evidence (local):
  python3 scripts/check-contributor-credit.py -> all credited on three surfaces
  scripts/sync-changelog.sh --check -> up to date
  scripts/release/check-versions.sh --range-audit-advisory -> OK
  check-feature-release-notes.sh <base> HEAD -> OK
  check-bundled-plugin-claims.py, check-ohos-deps.sh -> OK
  vitest lib/public-copy.test.ts lib/public-surface-contract.test.ts -> 18 passed

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Resolve the receipt cwd asynchronously before process creation and pass that
same resolved path to the process manager and OS. A command that retargets
the workspace symlink cannot change the reported starting directory.

Clear inherited DEEPSEEK_TOOL_EXECUTION_RECEIPT before overlaying the current
hook context in both foreground and background hook launches. Clarify that
bounded previews show retained tool output, which may already omit earlier
process output. Preserve the original PR and contributor attribution.

Validation on the host using the existing contributor target and hermetic
test HOME:
- cargo test --locked -p codewhale-tui --lib -- execution_receipt
  runtime_tool_completion runtime_shell_completion --test-threads=1:
  7 passed, 0 failed (including symlink-retarget and six actual hook children
  covering absent, oversized and current receipts in both launch modes).
- npm test: 636 passed, 0 failed.
- npm run check:web: passed.
- rustfmt and git diff --check: passed.

The child sandbox's earlier npm and web attempts failed on loopback EPERM
and api.github.com DNS; the host reruns above completed successfully.
No provider, deployment or release qualification is claimed.
Merge 86aa086 without rewriting the
original PR or receipt-repair checkpoint. Resolve the sole public facts
credit-array conflict by retaining @wuisabel-gif, @BX166 and @cenab once
each; preserve both sides' changelog and contributor entries.

Merged-tree validation on the host:
- npm test: 636 passed, 0 failed.
- npm run check:web: passed, including the production web build.
- git diff --check and public-facts JSON/credit uniqueness: passed.

The receipt repair checkpoint 69eebf8 passed 7 focused Rust tests before
this merge. Source inspection confirms the hook implementation is identical
and main's only shell implementation delta is its unrelated idle-PTY stale
guard; those tests were not rerun solely for merging. Hosted CI remains a
separate gate. No provider, deploy or release claim.
Resolve four changelog/contributor-credit conflicts by preserving both credit sets. The seven original hook implementation, test and documentation patches are unchanged after normalization, including the Runtime metadata exclusion.

Validation: npm test 670 passed, 0 failed; npm run check:web passed; version/credit checks passed. Root and independent source-preservation reviews passed. Child HTTP fixtures were blocked by loopback EPERM; the required root gate passed unrestricted. No Rust source conflict was edited and no new local Rust build was claimed.
@Hmbown
Hmbown merged commit b66b4ab into main Sep 29, 2026
34 checks passed
@Hmbown
Hmbown deleted the feat/6689-hook-execution-receipt branch September 29, 2026 14:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

hooks: export a post-admission execution receipt to tool_call_after

2 participants