Description
A tool call that fails is persisted without any tool result: the session item keeps
status: "failed" and metadata.tool_result_for: null, while a successful call carries
its own id in that field. While the same runtime process stays alive this is invisible —
the failure is still in memory and the model receives it normally. After the runtime is
restarted, the next request is rebuilt from the persisted transcript and contains a
function_call with no matching function_call_output, so the provider rejects the
whole request:
Responses API request failed: Invalid request (400):
No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008.
The thread is then permanently wedged: every following turn fails in about one second,
and no Runtime API call repairs it.
Steps to reproduce
- Start the runtime:
app-server --http --host 127.0.0.1 --port <port> --auth-token <token>
(home passed through APPDATA/LOCALAPPDATA/USERPROFILE, as the vendor launchers do)
POST /v1/threads with {"workspace": "<empty directory>"}
POST /v1/threads/<id>/turns with
{"prompt": "Read the file net-takogo-fajla-12345.txt in the working directory and summarise its contents."}
The model calls read, the file does not exist, the tool fails.
This turn completes normally.
- Restart
app-server on the same home.
POST /v1/threads/<id>/turns with {"prompt": "Answer with one word: alive."}
→ the turn fails with 400 No tool output found for tool call ….
A script that performs exactly these five steps is attached to this report; it prints
REPRODUCED on a machine where the bug is present (exit code 0). It talks to a real
provider — two turns, a fraction of a cent.
Expected behavior
A failed tool call must leave the transcript replayable. Either record a tool result for
it (is_error: true together with the failure text), or drop the call — or synthesise a
function_call_output — when the request is built. Restarting the runtime must not make
an existing thread unsendable.
Actual behavior
The persisted transcript contains a call with nothing answering it:
failed call: {"kind":"tool_call","status":"failed",
"metadata":{"tool_use_id":"call_00_xSP…|a372d3f5-…",
"tool_result_for":null}}
completed call: {"kind":"tool_call","status":"completed",
"metadata":{"tool_use_id":"call_00_kwb…|b3dfeb42-…",
"tool_result_for":"call_00_kwb…|b3dfeb42-…"}}
Within one runtime process nothing breaks; after a restart every turn fails. There is no
repair route:
POST /v1/threads/<id>/turns/<turn>/tool-calls/<call_id>/result accepts only
pending dynamic tool calls — for a failed call from an earlier turn it answers
404 {"message":"No pending dynamic tool call '<id>'"} (and the thread reports
pending_dynamic_tool_calls: []).
POST /v1/threads/<id>/fork copies the whole thread, orphan included; it ignores
unknown body fields and offers no truncation.
- Only
POST /v1/threads/<id>/undo helps, and it removes exactly one exchange per call
and returns a new thread.
Reproduced live on 2026-09-30, deepseek-flash:
turn 1 : completed
tool_call failed tool_result_for=None
tool_call completed tool_result_for='call_00_DVQd6AoxV6ia428T0Gwg8053|edaff0d0-…'
tool_call completed tool_result_for='call_00_cIVEAJs8WjwKBojU13Q30537|ab7657cf-…'
tool_call completed tool_result_for='call_01_4NofpBOZf2PERQj72T7p6773|13c6e540-…'
tool_call completed tool_result_for='call_00_WAq4k4cKMkEsg0y5SvWs4846|079c1917-…'
failed calls with no tool_result_for: 1
--- restarting the runtime on the same home
turn : turn_65d3d034 failed
error failed … (400): No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008
== REPRODUCED
Impact
Happens 100% of the time once a tool has failed and the app is restarted — after that the
thread cannot be used at all, and every new message is rejected in about a second.
Triggers seen so far, all of them ordinary: a code_execution call timing out after
120 s, a code_execution call failing input validation, and a read of a missing path.
The restart is not exotic either — it happens on every app restart, and our client
restarts the runtime whenever settings are saved.
The workaround costs the user history: N undo calls for a broken turn that sits N turns
from the end, each one leaving another thread behind (there is no delete).
Environment
-
OS: Windows 10 Pro, 10.0.19045 (AMD64)
-
codewhale version: codewhale 0.10.0 (1be1a703b975); runtime_api_version: 1.0
-
Install method: release binary (codew-windows-x64.exe), run as app-server --http
-
codewhale doctor summary (relevant lines only):
Version Information:
codewhale-tui: 0.10.0 (1be1a703b975)
API Connectivity:
· provider: deepseek
· base_url: https://api.deepseek.com
· model: deepseek-flash (resolved)
· strict_tool_mode: disabled
Tool Dependencies:
✓ Python: python → code_execution tool registered
✓ Node.js: present → js_execution tool registered
-
Model/provider: deepseek-flash via deepseek (https://api.deepseek.com)
-
Terminal app: none — the runtime is started headless by a browser client that talks to
/v1/* and renders the SSE stream (the same thread had also been used from the TUI
earlier that day)
-
Shell: PowerShell 5.1 on Windows
Logs, screenshots, or recordings
-
The reproduction script is attached to this report (steps 1–5 above, with the verdict).
-
Nothing comes from the provider side except the rejection itself, because the request is
refused before it reaches the model:
Responses API request failed: Invalid request (400):
No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008. (request_id: …)
-
Notes:
max_history: 100 and auto_compact: false in this config are irrelevant — the
reproduction runs on a 17-item thread.
- The same missing result is recorded when the provider is a local OpenAI-compatible
endpoint (that is where we first saw it); such an endpoint does not reject unpaired
tool calls, so the hole stayed invisible there.
strict_tool_mode: disabled here — would enabling it catch a tool call that has no
output before the request goes out?
- Side observation, not investigated: after
undo the runtime rewrote the session —
turn ids changed while the thread id stayed the same.
repro_orphan_tool_call.py
Description
A tool call that fails is persisted without any tool result: the session item keeps
status: "failed"andmetadata.tool_result_for: null, while a successful call carriesits own id in that field. While the same runtime process stays alive this is invisible —
the failure is still in memory and the model receives it normally. After the runtime is
restarted, the next request is rebuilt from the persisted transcript and contains a
function_callwith no matchingfunction_call_output, so the provider rejects thewhole request:
The thread is then permanently wedged: every following turn fails in about one second,
and no Runtime API call repairs it.
Steps to reproduce
app-server --http --host 127.0.0.1 --port <port> --auth-token <token>(home passed through
APPDATA/LOCALAPPDATA/USERPROFILE, as the vendor launchers do)POST /v1/threadswith{"workspace": "<empty directory>"}POST /v1/threads/<id>/turnswith{"prompt": "Read the file net-takogo-fajla-12345.txt in the working directory and summarise its contents."}The model calls
read, the file does not exist, the tool fails.This turn completes normally.
app-serveron the same home.POST /v1/threads/<id>/turnswith{"prompt": "Answer with one word: alive."}→ the turn fails with
400 No tool output found for tool call ….A script that performs exactly these five steps is attached to this report; it prints
REPRODUCEDon a machine where the bug is present (exit code 0). It talks to a realprovider — two turns, a fraction of a cent.
Expected behavior
A failed tool call must leave the transcript replayable. Either record a tool result for
it (
is_error: truetogether with the failure text), or drop the call — or synthesise afunction_call_output— when the request is built. Restarting the runtime must not makean existing thread unsendable.
Actual behavior
The persisted transcript contains a call with nothing answering it:
Within one runtime process nothing breaks; after a restart every turn fails. There is no
repair route:
POST /v1/threads/<id>/turns/<turn>/tool-calls/<call_id>/resultaccepts onlypending dynamic tool calls — for a failed call from an earlier turn it answers
404 {"message":"No pending dynamic tool call '<id>'"}(and the thread reportspending_dynamic_tool_calls: []).POST /v1/threads/<id>/forkcopies the whole thread, orphan included; it ignoresunknown body fields and offers no truncation.
POST /v1/threads/<id>/undohelps, and it removes exactly one exchange per calland returns a new thread.
Reproduced live on 2026-09-30,
deepseek-flash:Impact
Happens 100% of the time once a tool has failed and the app is restarted — after that the
thread cannot be used at all, and every new message is rejected in about a second.
Triggers seen so far, all of them ordinary: a
code_executioncall timing out after120 s, a
code_executioncall failing input validation, and areadof a missing path.The restart is not exotic either — it happens on every app restart, and our client
restarts the runtime whenever settings are saved.
The workaround costs the user history: N
undocalls for a broken turn that sits N turnsfrom the end, each one leaving another thread behind (there is no delete).
Environment
OS: Windows 10 Pro, 10.0.19045 (AMD64)
codewhale version:
codewhale 0.10.0 (1be1a703b975);runtime_api_version: 1.0Install method: release binary (
codew-windows-x64.exe), run asapp-server --httpcodewhale doctorsummary (relevant lines only):Model/provider:
deepseek-flashviadeepseek(https://api.deepseek.com)Terminal app: none — the runtime is started headless by a browser client that talks to
/v1/*and renders the SSE stream (the same thread had also been used from the TUIearlier that day)
Shell: PowerShell 5.1 on Windows
Logs, screenshots, or recordings
The reproduction script is attached to this report (steps 1–5 above, with the verdict).
Nothing comes from the provider side except the rejection itself, because the request is
refused before it reaches the model:
Notes:
max_history: 100andauto_compact: falsein this config are irrelevant — thereproduction runs on a 17-item thread.
endpoint (that is where we first saw it); such an endpoint does not reject unpaired
tool calls, so the hole stayed invisible there.
strict_tool_mode: disabledhere — would enabling it catch a tool call that has nooutput before the request goes out?
undothe runtime rewrote the session —turn ids changed while the thread id stayed the same.
repro_orphan_tool_call.py