Skip to content

A failed tool call leaves no tool output, so the thread becomes unsendable after a runtime restart. #6803

Description

@Denis-VG

Description

A tool call that fails is persisted without any tool result: the session item keeps
status: "failed" and metadata.tool_result_for: null, while a successful call carries
its own id in that field. While the same runtime process stays alive this is invisible —
the failure is still in memory and the model receives it normally. After the runtime is
restarted, the next request is rebuilt from the persisted transcript and contains a
function_call with no matching function_call_output, so the provider rejects the
whole request:

Responses API request failed: Invalid request (400):
No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008.

The thread is then permanently wedged: every following turn fails in about one second,
and no Runtime API call repairs it.

Steps to reproduce

  1. Start the runtime:
    app-server --http --host 127.0.0.1 --port <port> --auth-token <token>
    (home passed through APPDATA/LOCALAPPDATA/USERPROFILE, as the vendor launchers do)
  2. POST /v1/threads with {"workspace": "<empty directory>"}
  3. POST /v1/threads/<id>/turns with
    {"prompt": "Read the file net-takogo-fajla-12345.txt in the working directory and summarise its contents."}
    The model calls read, the file does not exist, the tool fails.
    This turn completes normally.
  4. Restart app-server on the same home.
  5. POST /v1/threads/<id>/turns with {"prompt": "Answer with one word: alive."}
    → the turn fails with 400 No tool output found for tool call ….

A script that performs exactly these five steps is attached to this report; it prints
REPRODUCED on a machine where the bug is present (exit code 0). It talks to a real
provider — two turns, a fraction of a cent.

Expected behavior

A failed tool call must leave the transcript replayable. Either record a tool result for
it (is_error: true together with the failure text), or drop the call — or synthesise a
function_call_output — when the request is built. Restarting the runtime must not make
an existing thread unsendable.

Actual behavior

The persisted transcript contains a call with nothing answering it:

failed call:    {"kind":"tool_call","status":"failed",
                 "metadata":{"tool_use_id":"call_00_xSP…|a372d3f5-…",
                             "tool_result_for":null}}

completed call: {"kind":"tool_call","status":"completed",
                 "metadata":{"tool_use_id":"call_00_kwb…|b3dfeb42-…",
                             "tool_result_for":"call_00_kwb…|b3dfeb42-…"}}

Within one runtime process nothing breaks; after a restart every turn fails. There is no
repair route:

  • POST /v1/threads/<id>/turns/<turn>/tool-calls/<call_id>/result accepts only
    pending dynamic tool calls — for a failed call from an earlier turn it answers
    404 {"message":"No pending dynamic tool call '<id>'"} (and the thread reports
    pending_dynamic_tool_calls: []).
  • POST /v1/threads/<id>/fork copies the whole thread, orphan included; it ignores
    unknown body fields and offers no truncation.
  • Only POST /v1/threads/<id>/undo helps, and it removes exactly one exchange per call
    and returns a new thread.

Reproduced live on 2026-09-30, deepseek-flash:

turn 1 : completed
  tool_call failed     tool_result_for=None
  tool_call completed  tool_result_for='call_00_DVQd6AoxV6ia428T0Gwg8053|edaff0d0-…'
  tool_call completed  tool_result_for='call_00_cIVEAJs8WjwKBojU13Q30537|ab7657cf-…'
  tool_call completed  tool_result_for='call_01_4NofpBOZf2PERQj72T7p6773|13c6e540-…'
  tool_call completed  tool_result_for='call_00_WAq4k4cKMkEsg0y5SvWs4846|079c1917-…'
  failed calls with no tool_result_for: 1
--- restarting the runtime on the same home
turn   : turn_65d3d034 failed
  error failed  … (400): No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008
== REPRODUCED

Impact

Happens 100% of the time once a tool has failed and the app is restarted — after that the
thread cannot be used at all, and every new message is rejected in about a second.

Triggers seen so far, all of them ordinary: a code_execution call timing out after
120 s, a code_execution call failing input validation, and a read of a missing path.
The restart is not exotic either — it happens on every app restart, and our client
restarts the runtime whenever settings are saved.

The workaround costs the user history: N undo calls for a broken turn that sits N turns
from the end, each one leaving another thread behind (there is no delete).

Environment

  • OS: Windows 10 Pro, 10.0.19045 (AMD64)

  • codewhale version: codewhale 0.10.0 (1be1a703b975); runtime_api_version: 1.0

  • Install method: release binary (codew-windows-x64.exe), run as app-server --http

  • codewhale doctor summary (relevant lines only):

    Version Information:
      codewhale-tui: 0.10.0 (1be1a703b975)
    
    API Connectivity:
      · provider: deepseek
      · base_url: https://api.deepseek.com
      · model: deepseek-flash (resolved)
      · strict_tool_mode: disabled
    
    Tool Dependencies:
      ✓ Python: python → code_execution tool registered
      ✓ Node.js: present → js_execution tool registered
    
  • Model/provider: deepseek-flash via deepseek (https://api.deepseek.com)

  • Terminal app: none — the runtime is started headless by a browser client that talks to
    /v1/* and renders the SSE stream (the same thread had also been used from the TUI
    earlier that day)

  • Shell: PowerShell 5.1 on Windows

Logs, screenshots, or recordings

  • The reproduction script is attached to this report (steps 1–5 above, with the verdict).

  • Nothing comes from the provider side except the rejection itself, because the request is
    refused before it reaches the model:

    Responses API request failed: Invalid request (400):
    No tool output found for tool call call_00_iwWH8V1nyzeeIJbRbRlv5008. (request_id: …)
    
  • Notes:

    • max_history: 100 and auto_compact: false in this config are irrelevant — the
      reproduction runs on a 17-item thread.
    • The same missing result is recorded when the provider is a local OpenAI-compatible
      endpoint (that is where we first saw it); such an endpoint does not reject unpaired
      tool calls, so the hole stayed invisible there.
    • strict_tool_mode: disabled here — would enabling it catch a tool call that has no
      output before the request goes out?
    • Side observation, not investigated: after undo the runtime rewrote the session —
      turn ids changed while the thread id stayed the same.

repro_orphan_tool_call.py

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingneeds-triageNew external report awaiting maintainer triage; repro, logs and version output help

    Projects

    • Status
      Backlog

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions