Conversation
…g long executions Root cause: every MCP request shared one finite wall-clock cap sized for cheap metadata calls. The engine-side stdio proxy answered tools/call under the generic 120s request budget, and the TUI pool defaulted its execute timeout to 60s; a legitimate tool run that takes minutes — a build, a test suite, a long script driven through an MCP server — was killed mid-flight and the model got "timed out" for healthy work, then retried and compounded the cost. Worse, raising the documented `execute_timeout` knob did not actually govern: the connection's inner response read wait stayed at `read_timeout`, fired first, marked the connection Disconnected, and silently cut the call off at the read knob. Mechanism: - crates/mcp stdio proxy: tools/call now uses a dedicated CALL_TOOL_TIMEOUT; every other request keeps the generic 120s REQUEST_TIMEOUT. - TUI pool: default execute timeout 60s -> 1800s. The per-server and global `execute_timeout` config fields remain the override path (`effective_execute_timeout`), so the constants are documented defaults only. - TUI connection: the per-request inner read wait is widened to max(read_timeout, that request's own budget), so a server that is silent for the whole execution of a long tools/call is not declared dead mid-call. Requests whose budget does not exceed the read knob (resources/read, discovery) still fail at the configured read budget. - TUI HTTP transports: the client's total request ceiling is now max(read, execute) so a raised execute_timeout governs HTTP servers too; the read knob itself stays intact for the connection-level waits. Numbers: 1800s (30 minutes) covers long-but-bounded tool workloads (builds, test suites, remote jobs) while staying a real bound — a wedged server still cannot hang a consumer forever. The stdio constant matches the TUI pool default so both surfaces behave the same for one server. Coordination: PR Hmbown#6711 touches stream-open budgets in the engine; this change deliberately does not touch stream-open or retry logic — it only re-budgets per-request MCP calls. Signed-off-by: asto18089 <asto18089@126.com>
asto18089
force-pushed
the
upstream/mcp-tools-call-budget
branch
from
September 29, 2026 11:45
be27103 to
3ffbb90
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
MCP tool calls were killed by two stacked short budgets:
crates/mcp's genericREQUEST_TIMEOUT(120s) applied to all requests includingtools/call, and the TUI pool'sdefault_execute_timeout(60s). A legitimate tool call that runs minutes — builds, test suites, scrapes, remote jobs — was killed mid-flight, and the model's retry compounded the cost. The read-wait also undercut both:recvbounded every response wait by the 120s read knob, so even a raisedexecute_timeoutfired first at the inner read and marked the connection Disconnected.Changes:
crates/mcp:tools/callgets a dedicatedCALL_TOOL_TIMEOUT(30 min, matching the TUI pool default); every other request (initialize, lists, resources) keeps the 120s generic budget.default_execute_timeout60s → 30 min; per-server/globalexecute_timeoutoverrides still win via the existingeffective_execute_timeoutseam.max(read_timeout, request budget), so the read knob can no longer silently defeat a raised execute budget; requests whose budget doesn't exceed the read knob keep failing fast at it.max(read, execute)so HTTP servers get the same semantics.Known trade-offs, disclosed for review:
prompts/getroutes througheffective_execute_timeout, so it inherits the new default (comment in code notes this; the override knob is the intended way to keep it tight).tools/callnow blocks sibling-server calls/status for up to 30 min (previously ≤60s) unless the turn is interrupted — same shape as before, longer tail; the per-serverexecute_timeoutoverride is the escape hatch. Happy to explore hold-lock-only-for-dispatch as a follow-up.Coordination: upstream #6711 reworks engine stream-open budgets — different surface, no file overlap.
Adapted from the Pinvou fork's timeout audit (Pinvou/CodeWhale
d349f2537).Testing
cargo fmt --all -- --checkcargo clippy -p codewhale-mcp -p codewhale-tui --all-features --locked(clean under the CI allow list)cargo test -p codewhale-mcp --lib— 83 passed;cargo test -p codewhale-tui --lib mcp::— 251 passed. New pins: a slowtools/calloutlives an injected generic budget whileresources/readstill fails at it; a tool response after the read knob still completes within the execute budget (red without the widening); requests under the read knob still fail at it; the HTTP ceiling covers the execute budgetChecklist
docs/MCP.md+ zh minimal example now shows the new default)CHANGELOG.mdchanges