Skip to content

Stall recovery is UI-side only: the engine keeps the wedged turn, the next send is refused for 60 s, and the app stops accepting input #6800

Description

@7jrxt42BxFZo4iAnN4CX

Goal / Why

Codewhale detects a wedged turn in the UI and "recovers" from it — but the recovery only resets UI
turn state. The engine keeps the same turn in flight, so the app and the engine disagree: the next
message cannot be admitted (60 s dispatch bound, then refused), the transcript stops accepting input
(Esc and Ctrl+C — the remedy the app itself prints — are dead), no turn outcome is persisted, the
idle inhibitor stays held, and only SIGKILL from another terminal ends it. The app meanwhile reports
recovery through its own toast.

Observed (one session; records on disk)

Session 19e2085f-bde4-45af-88cb-9401acb33568, operate mode, command-code / deepseek/deepseek-v4.1-flash.
All times UTC.

  • 14:18:53 session file last write. The parent had called agent {"action":"wait","until":"completion"};
    the wait timed out at 30 s with the child still running, then the child's subagent_completion
    runtime event arrived — the last message persisted.
  • 14:30:19 the UI's tool-hang branch fires → crashes/20260930T143019.553Z-turn-stall-ui.log
    (while a tool ran with no progress, 600 s, turn 0ab75777-…) and recover_stalled_runtime_turn
    runs. Its own toast says: "Tool stalled with no progress for 10m — recovered; the command may still
    be running in the background. Use exec_shell_cancel or retry." A recovery checkpoint is written in
    the same window (~/.codewhale/sessions/checkpoints/19e2085f-….json).
  • 14:31:21 / :30 / :42 snapshot prune rounds back to back (context only — see Out of scope).
  • 14:31:49 the operator's next message is not admitted for 60 s (DISPATCH_TASK_BOUND) →
    crashes/20260930T143149.561Z-turn-stall-ui.log (while dispatching the message (route planning / engine admission)) and "Message dispatch stalled for 60s before the engine accepted it; your
    message was restored. Press Esc to cancel the running turn, then retry." — the engine still owned
    the turn, i.e. the 14:30 recovery had not touched it.
  • 14:38:09 the engine's own heartbeat reports the same turn stalled in Streaming
    (crates/tui/src/core/engine/turn_heartbeat.rs:62), 334 s vs bound 330 s:
    crashes/20260930T143809.565Z-turn-stall-engine.log.
  • Process state at that point: 0.0 % CPU over a 3 s sample, 30 threads, all S, no child processes;
    systemd-inhibit --what=idle --why="Codewhale turn in flight" --mode=block cat still held.
  • The session file was never written again (still 185 messages, turn_outcomes: [], hours later); the
    in-flight turn had no outcome, and after a restart the operator could not read the sub-agent's
    output either (tool=handle_read … no payload found for handle agent:agent_1b13afd6/transcript).

The same shape recurred the same day (crashes/20260930T020236.312Z-turn-stall-ui.log, tool phase
600 s, turn 719b5c03-…). Not deterministically reproducible yet; it appears with a wedged provider
wait together with concurrent workspace churn.

Scope / Plan

  1. crates/tui/src/tui/ui/session_state.rs:447 — recover_stalled_runtime_turn must cancel the
    engine's turn (the same cancellation the operator's cancel key issues) before clearing UI state.
    If the engine cannot be reached, the UI must not report recovery: it must offer a force-end that
    ends the engine's turn locally (or restarts the engine task) and records the outcome.
  2. crates/tui/src/tui/ui/session_state.rs:334-392 — the two branches that call it keep firing, but
    the recovery must be all-or-nothing: engine cancelled + outcome persisted + idle inhibitor
    released.
  3. crates/tui/src/tui/ui/dispatch.rs:725/736/774/794 — the dispatch wait must not block input for
    DISPATCH_TASK_BOUND; the operator must be able to type and cancel while a send is pending. The
    timeout path itself is fine (the message is restored) — the blocking is not.
  4. crates/tui/src/core/engine/turn_heartbeat.rs:386-411 — the heartbeat already records and reports
    (:341 detect, :426 the existing "once per episode" contract); keep that, and make sure the
    escalation contract (report → turn ends) is what actually closes the turn, so a phase whose inner
    bound failed (:35-37) cannot outlive the report.
  5. Persist a turn outcome on any forced end (interrupted/failed + reason) and release the turn's
    idle inhibitor, so a restart does not silently continue from a half-finished turn.

Key files

crates/tui/src/tui/ui/session_state.rs
crates/tui/src/tui/ui/dispatch.rs
crates/tui/src/core/engine/turn_heartbeat.rs
crates/tui/src/core/engine/turn_loop.rs

Acceptance criteria

  • After a stall recovery a new message is admitted immediately — no 60 s refusal, no "engine
    still owns the turn" state.
  • The transcript accepts input during and after recovery; Esc and Ctrl+C always cancel.
  • A forced end persists a turn outcome (interrupted/failed + the stall reason) and releases
    the idle inhibitor.
  • A sub-agent completion already delivered into the transcript is not silently unreachable after
    a forced end.
  • Recovery never fires for a turn that is making progress.

Verification

cargo check -p codewhale-tui
cargo test -p codewhale-tui -- session_state
cargo test -p codewhale-tui -- dispatch
cargo test -p codewhale-tui -- turn_heartbeat
cargo clippy -p codewhale-tui -- -D warnings

New tests: recovery cancels the engine turn (assert the engine no longer owns a turn and the next
send is admitted); dispatch keeps the event loop responsive while a send is pending; an outcome is
persisted on recovery. The existing heartbeat contract test (turn_heartbeat.rs:426) must keep
passing.

Out of scope

Environment

  • Arch Linux, kernel 7.2.7-arch1-1, x86_64; rustc/cargo 1.97.1
  • codewhale 0.10.1 (dev), source build, binary built 2026-09-29 16:59 +0300 (= 13:59Z)
  • session mode operate; provider command-code (deepseek/deepseek-v4.1-flash), OpenAI-compatible
  • [tui] stream_max_resumes = 10, stream_max_transparent_retries = 10, stream_max_errors = 10,
    turn_wall_clock_secs = 86400 (no global turn bound); stream_chunk_timeout_secs unset — the
    effective inner bound in that session was 330 s (300 s + the 30 s STALL_BOUND_GRACE)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageNew external report awaiting maintainer triage; repro, logs and version output help

    Projects

    • Status
      Backlog

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions