feat(hooks): export a post-admission execution receipt to tool_call_after (#6689) - #6713
Merged
Merged
Conversation
…fter (#6689) tool_call_after hooks could not tell what a shell tool actually ran. They got the result text, the success flag, the exit code and the status, but not the command or its directory. A tool_call_before rewrite can change the before-hook input they would otherwise re-derive those from. DEEPSEEK_TOOL_EXECUTION_RECEIPT now carries schema-1 JSON with these fields: command, cwd, state (completed/interrupted), scope, exit_code (observed or null), bounded stdout/stderr previews, truncation flags and output_kind. - Identity reuses the spawn record. execute_foreground_via_background reads command and working_dir from the BackgroundShell record the process manager created, under the same lock that marks the run as foreground. - HookContext::with_tool_outcome reads metadata.execution_receipt from Ok and Err results. That one seam covers the TUI and Runtime API completion hooks, and on_error for a failed shell call. Lowercase bash failures carry the receipt too. - Exact or absent. An identity over 8 KiB, containing NUL, or with a relative or non-UTF-8 directory gets no receipt. Previews keep both ends and halve until the serialized JSON is at most 32 KiB. The hook boundary drops a non-schema-1 or oversized receipt instead of truncating it. - Scope is settled, pipe-backed, unsandboxed, local foreground runs only. No receipt for background, moved-to-/jobs, PTY/interactive, OS sandbox, external backend, read-only argv hardening, Windows, or pre-exec refusals. - Documented in docs/HOOKS.md and docs/zh_hans/HOOKS.md. The contract and test design follow Isabel Wu's reference branch wuisabel-gif/CodeWhale@feat/tool-call-after-execution-receipt. This version is re-implemented on the current with_tool_outcome seam and on main's process-manager record. Evidence (macOS, CARGO_TARGET_DIR outside the repo): cargo test -p codewhale-tui --lib -- execution_receipt -> test result: ok. 3 passed; 0 failed; 0 ignored; 13816 filtered out same run with the metadata read and identity capture disabled -> test result: FAILED. 1 passed; 2 failed (both new behavior tests fail) cargo test -p codewhale-tui --lib -- hooks:: tools::shell:: tool_routing -> test result: ok. 340 passed; 0 failed; 1 ignored cargo test -p codewhale-tui --lib -- runtime_tool_completion_fires_after_and_error_hooks runtime_shell_completion_delivers_exit_code_and_status_to_hooks -> test result: ok. 2 passed; 0 failed cargo clippy -p codewhale-tui --lib --tests --locked -- -D warnings (CI allow-list) -> Finished, 0 warnings cargo fmt --all -- --check -> clean Co-authored-by: Isabel Wu <231155141+wuisabel-gif@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Comment on lines
+5025
to
+5026
| let (stdout, stdout_clipped) = execution_receipt_preview(&result.stdout, budget); | ||
| let (stderr, stderr_clipped) = execution_receipt_preview(&result.stderr, budget); |
Contributor
There was a problem hiding this comment.
🔍 Receipt previews inherit earlier output clipping
For long runs, shell_execution_receipt previews the already-bounded ShellResult, not the raw streams. The lowercase bash accumulator can retain only a tail, so its receipt cannot show the original first bytes. The documented preview contract may need to distinguish tool output from raw process output.
Was this helpful? React with 👍 or 👎 to provide feedback.
Review follow-ups on the tool_call_after execution receipt:
- cwd has one spelling. An explicit cwd reached the spawn already
canonicalized by ToolContext::resolve_path, while the default workspace
was passed as the session opened it, so a workspace opened through a
symlink reported two paths for one directory. The receipt now
canonicalizes the recorded directory with tokio::fs::canonicalize (off
the runtime thread); a directory that no longer resolves gets no receipt.
The test opens the workspace through a symlink and compares exact strings
instead of re-canonicalizing them.
- A wait error is not an interruption. BackgroundShell::poll marks a failed
try_wait (wait_failed); such a run gets no receipt instead of state
"interrupted" for a process that may have exited normally.
- PowerShell is excluded on every platform. $SHELL=pwsh on Unix makes the
dispatcher wrap the source or run it from a temp -File, so the receipt's
command would not be what ran.
- Hook-only. The receipt is built only while a tool_call_after or on_error
hook is registered, and the Runtime API strips it before persisting and
emitting the item, which already carries the output.
- docs/HOOKS.md and docs/zh_hans/HOOKS.md describe all four.
Evidence (macOS, CARGO_TARGET_DIR outside the repo):
cargo test -p codewhale-tui --lib --locked -- execution_receipt
-> test result: ok. 3 passed; 0 failed; 0 ignored; 0 measured; 13816 filtered out
same test with the canonicalize step removed
-> test result: FAILED. 0 passed; 1 failed (left: .../link, right: /private/var/.../real)
cargo test -p codewhale-tui --lib --locked -- hooks:: tools::shell:: tool_routing runtime_tool_completion runtime_shell_completion
-> test result: ok. 341 passed; 0 failed; 1 ignored; 0 measured; 13477 filtered out
cargo clippy -p codewhale-tui --lib --tests --locked -- -D warnings (CI allow-list)
-> Finished, 0 warnings
cargo fmt --all -- --check -> clean
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…6689) b965cd7 carries a Co-authored-by trailer for Isabel Wu, whose reference branch set the tool_call_after execution-receipt contract and test design. check-contributor-credit.py (Version drift job) failed because the handle was not in web/lib/release-credits.ts. Add her to the v0.10.1 band on every credit surface so the parity tests stay exact: RELEASE_CONTRIBUTORS, the 0.10.1 CHANGELOG Contributors block (and the synced crates/tui copy), the docs/CONTRIBUTORS.md v0.10.1 band, and requiredCandidateCredits. Evidence (local): python3 scripts/check-contributor-credit.py -> all credited on three surfaces scripts/sync-changelog.sh --check -> up to date scripts/release/check-versions.sh --range-audit-advisory -> OK check-feature-release-notes.sh <base> HEAD -> OK check-bundled-plugin-claims.py, check-ohos-deps.sh -> OK vitest lib/public-copy.test.ts lib/public-surface-contract.test.ts -> 18 passed Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Resolve the receipt cwd asynchronously before process creation and pass that same resolved path to the process manager and OS. A command that retargets the workspace symlink cannot change the reported starting directory. Clear inherited DEEPSEEK_TOOL_EXECUTION_RECEIPT before overlaying the current hook context in both foreground and background hook launches. Clarify that bounded previews show retained tool output, which may already omit earlier process output. Preserve the original PR and contributor attribution. Validation on the host using the existing contributor target and hermetic test HOME: - cargo test --locked -p codewhale-tui --lib -- execution_receipt runtime_tool_completion runtime_shell_completion --test-threads=1: 7 passed, 0 failed (including symlink-retarget and six actual hook children covering absent, oversized and current receipts in both launch modes). - npm test: 636 passed, 0 failed. - npm run check:web: passed. - rustfmt and git diff --check: passed. The child sandbox's earlier npm and web attempts failed on loopback EPERM and api.github.com DNS; the host reruns above completed successfully. No provider, deployment or release qualification is claimed.
Merge 86aa086 without rewriting the original PR or receipt-repair checkpoint. Resolve the sole public facts credit-array conflict by retaining @wuisabel-gif, @BX166 and @cenab once each; preserve both sides' changelog and contributor entries. Merged-tree validation on the host: - npm test: 636 passed, 0 failed. - npm run check:web: passed, including the production web build. - git diff --check and public-facts JSON/credit uniqueness: passed. The receipt repair checkpoint 69eebf8 passed 7 focused Rust tests before this merge. Source inspection confirms the hook implementation is identical and main's only shell implementation delta is its unrelated idle-PTY stale guard; those tests were not rerun solely for merging. Hosted CI remains a separate gate. No provider, deploy or release claim.
Resolve four changelog/contributor-credit conflicts by preserving both credit sets. The seven original hook implementation, test and documentation patches are unchanged after normalization, including the Runtime metadata exclusion. Validation: npm test 670 passed, 0 failed; npm run check:web passed; version/credit checks passed. Root and independent source-preservation reviews passed. Child HTTP fixtures were blocked by loopback EPERM; the required root gate passed unrestricted. No Rust source conflict was edited and no new local Rust build was claimed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #6689. Refs #6582; that issue remains open because this uses the proposed environment variable rather than a stdin JSON envelope, with
stateandoutput_kindfield names.After an admitted foreground shell command runs,
tool_call_afterand failed-callon_errorhooks receive a schema-1 JSONDEEPSEEK_TOOL_EXECUTION_RECEIPT. It carries the command from the actual spawn record, the directory passed to the process, completion state, observed exit code, and bounded stdout/stderr previews. Both the TUI and Runtime API use the existingHookContext::with_tool_outcomeboundary; the Runtime API strips this hook-only metadata before persisting or emitting the tool item.The directory is resolved asynchronously before spawning and that same path supplies process and receipt identity. A command that retargets a workspace symlink cannot change the reported starting directory. Both foreground and background hook launchers clear an inherited receipt before overlaying the current call's context, so an absent or oversized current receipt cannot expose another call's receipt.
Receipts remain exact or absent: oversized or invalid identities, unresolved directories, failed waits, PTY/interactive, background/detached, sandboxed/external-backend, read-only argv, PowerShell and Windows paths do not claim an exact receipt. Preview limits apply to the tool's retained output; earlier process output may already have been dropped by its bounded capture. A too-large or unsupported-schema document is omitted intact rather than truncated. Hooks gain observation only, with no new admission or execution authority.
The contract and test design follow @wuisabel-gif's reference branch,
wuisabel-gif/CodeWhale@feat/tool-call-after-execution-receipt. The original implementation adapts that design to the current shared hook boundary and process-manager record, includes failed lowercasebashcalls, and preserves Isabel Wu's co-author attribution and release credits. Documentation is updated in English and Simplified Chinese.Validation:
codewhale-tuitests at repair checkpoint69eebf8ae: 7 passed, 0 failed. These execute a symlink-retargeting shell command and six real hook children, covering absent, oversized and current receipts in both launch modes.47a014c6isda1212d39ba365407970c0579e14b7d036c43a05.npm test: 670 passed, 0 failed (68 CLI, 16 runtime SDK, 50 extension host, 536 web).npm run check:web: passed, including production web build.The first child npm attempt could not bind its local HTTP fixtures; the root retry passed. Hosted CI on the latest head remains required. Local fixtures do not establish provider, native, deployment or release qualification.