Goal / Why
Codewhale detects a wedged turn in the UI and "recovers" from it — but the recovery only resets UI
turn state. The engine keeps the same turn in flight, so the app and the engine disagree: the next
message cannot be admitted (60 s dispatch bound, then refused), the transcript stops accepting input
(Esc and Ctrl+C — the remedy the app itself prints — are dead), no turn outcome is persisted, the
idle inhibitor stays held, and only SIGKILL from another terminal ends it. The app meanwhile reports
recovery through its own toast.
Observed (one session; records on disk)
Session 19e2085f-bde4-45af-88cb-9401acb33568, operate mode, command-code / deepseek/deepseek-v4.1-flash.
All times UTC.
- 14:18:53 session file last write. The parent had called
agent {"action":"wait","until":"completion"};
the wait timed out at 30 s with the child still running, then the child's subagent_completion
runtime event arrived — the last message persisted.
- 14:30:19 the UI's tool-hang branch fires →
crashes/20260930T143019.553Z-turn-stall-ui.log
(while a tool ran with no progress, 600 s, turn 0ab75777-…) and recover_stalled_runtime_turn
runs. Its own toast says: "Tool stalled with no progress for 10m — recovered; the command may still
be running in the background. Use exec_shell_cancel or retry." A recovery checkpoint is written in
the same window (~/.codewhale/sessions/checkpoints/19e2085f-….json).
- 14:31:21 / :30 / :42 snapshot prune rounds back to back (context only — see Out of scope).
- 14:31:49 the operator's next message is not admitted for 60 s (
DISPATCH_TASK_BOUND) →
crashes/20260930T143149.561Z-turn-stall-ui.log (while dispatching the message (route planning / engine admission)) and "Message dispatch stalled for 60s before the engine accepted it; your
message was restored. Press Esc to cancel the running turn, then retry." — the engine still owned
the turn, i.e. the 14:30 recovery had not touched it.
- 14:38:09 the engine's own heartbeat reports the same turn stalled in
Streaming
(crates/tui/src/core/engine/turn_heartbeat.rs:62), 334 s vs bound 330 s:
crashes/20260930T143809.565Z-turn-stall-engine.log.
- Process state at that point: 0.0 % CPU over a 3 s sample, 30 threads, all
S, no child processes;
systemd-inhibit --what=idle --why="Codewhale turn in flight" --mode=block cat still held.
- The session file was never written again (still 185 messages,
turn_outcomes: [], hours later); the
in-flight turn had no outcome, and after a restart the operator could not read the sub-agent's
output either (tool=handle_read … no payload found for handle agent:agent_1b13afd6/transcript).
The same shape recurred the same day (crashes/20260930T020236.312Z-turn-stall-ui.log, tool phase
600 s, turn 719b5c03-…). Not deterministically reproducible yet; it appears with a wedged provider
wait together with concurrent workspace churn.
Scope / Plan
crates/tui/src/tui/ui/session_state.rs:447 — recover_stalled_runtime_turn must cancel the
engine's turn (the same cancellation the operator's cancel key issues) before clearing UI state.
If the engine cannot be reached, the UI must not report recovery: it must offer a force-end that
ends the engine's turn locally (or restarts the engine task) and records the outcome.
crates/tui/src/tui/ui/session_state.rs:334-392 — the two branches that call it keep firing, but
the recovery must be all-or-nothing: engine cancelled + outcome persisted + idle inhibitor
released.
crates/tui/src/tui/ui/dispatch.rs:725/736/774/794 — the dispatch wait must not block input for
DISPATCH_TASK_BOUND; the operator must be able to type and cancel while a send is pending. The
timeout path itself is fine (the message is restored) — the blocking is not.
crates/tui/src/core/engine/turn_heartbeat.rs:386-411 — the heartbeat already records and reports
(:341 detect, :426 the existing "once per episode" contract); keep that, and make sure the
escalation contract (report → turn ends) is what actually closes the turn, so a phase whose inner
bound failed (:35-37) cannot outlive the report.
- Persist a turn outcome on any forced end (
interrupted/failed + reason) and release the turn's
idle inhibitor, so a restart does not silently continue from a half-finished turn.
Key files
crates/tui/src/tui/ui/session_state.rs
crates/tui/src/tui/ui/dispatch.rs
crates/tui/src/core/engine/turn_heartbeat.rs
crates/tui/src/core/engine/turn_loop.rs
Acceptance criteria
Verification
cargo check -p codewhale-tui
cargo test -p codewhale-tui -- session_state
cargo test -p codewhale-tui -- dispatch
cargo test -p codewhale-tui -- turn_heartbeat
cargo clippy -p codewhale-tui -- -D warnings
New tests: recovery cancels the engine turn (assert the engine no longer owns a turn and the next
send is admitted); dispatch keeps the event loop responsive while a send is pending; an outcome is
persisted on recovery. The existing heartbeat contract test (turn_heartbeat.rs:426) must keep
passing.
Out of scope
Environment
- Arch Linux, kernel 7.2.7-arch1-1, x86_64; rustc/cargo 1.97.1
- codewhale 0.10.1 (dev), source build, binary built 2026-09-29 16:59 +0300 (= 13:59Z)
- session mode
operate; provider command-code (deepseek/deepseek-v4.1-flash), OpenAI-compatible
[tui] stream_max_resumes = 10, stream_max_transparent_retries = 10, stream_max_errors = 10,
turn_wall_clock_secs = 86400 (no global turn bound); stream_chunk_timeout_secs unset — the
effective inner bound in that session was 330 s (300 s + the 30 s STALL_BOUND_GRACE)
Goal / Why
Codewhale detects a wedged turn in the UI and "recovers" from it — but the recovery only resets UI
turn state. The engine keeps the same turn in flight, so the app and the engine disagree: the next
message cannot be admitted (60 s dispatch bound, then refused), the transcript stops accepting input
(Esc and Ctrl+C — the remedy the app itself prints — are dead), no turn outcome is persisted, the
idle inhibitor stays held, and only SIGKILL from another terminal ends it. The app meanwhile reports
recovery through its own toast.
Observed (one session; records on disk)
Session
19e2085f-bde4-45af-88cb-9401acb33568, operate mode,command-code/deepseek/deepseek-v4.1-flash.All times UTC.
agent {"action":"wait","until":"completion"};the wait timed out at 30 s with the child still running, then the child's
subagent_completionruntime event arrived — the last message persisted.
crashes/20260930T143019.553Z-turn-stall-ui.log(
while a tool ran with no progress, 600 s, turn0ab75777-…) andrecover_stalled_runtime_turnruns. Its own toast says: "Tool stalled with no progress for 10m — recovered; the command may still
be running in the background. Use exec_shell_cancel or retry." A recovery checkpoint is written in
the same window (
~/.codewhale/sessions/checkpoints/19e2085f-….json).DISPATCH_TASK_BOUND) →crashes/20260930T143149.561Z-turn-stall-ui.log(while dispatching the message (route planning / engine admission)) and "Message dispatch stalled for 60s before the engine accepted it; yourmessage was restored. Press Esc to cancel the running turn, then retry." — the engine still owned
the turn, i.e. the 14:30 recovery had not touched it.
Streaming(
crates/tui/src/core/engine/turn_heartbeat.rs:62), 334 s vs bound 330 s:crashes/20260930T143809.565Z-turn-stall-engine.log.S, no child processes;systemd-inhibit --what=idle --why="Codewhale turn in flight" --mode=block catstill held.turn_outcomes: [], hours later); thein-flight turn had no outcome, and after a restart the operator could not read the sub-agent's
output either (
tool=handle_read … no payload found for handle agent:agent_1b13afd6/transcript).The same shape recurred the same day (
crashes/20260930T020236.312Z-turn-stall-ui.log, tool phase600 s, turn
719b5c03-…). Not deterministically reproducible yet; it appears with a wedged providerwait together with concurrent workspace churn.
Scope / Plan
crates/tui/src/tui/ui/session_state.rs:447—recover_stalled_runtime_turnmust cancel theengine's turn (the same cancellation the operator's cancel key issues) before clearing UI state.
If the engine cannot be reached, the UI must not report recovery: it must offer a force-end that
ends the engine's turn locally (or restarts the engine task) and records the outcome.
crates/tui/src/tui/ui/session_state.rs:334-392— the two branches that call it keep firing, butthe recovery must be all-or-nothing: engine cancelled + outcome persisted + idle inhibitor
released.
crates/tui/src/tui/ui/dispatch.rs:725/736/774/794— the dispatch wait must not block input forDISPATCH_TASK_BOUND; the operator must be able to type and cancel while a send is pending. Thetimeout path itself is fine (the message is restored) — the blocking is not.
crates/tui/src/core/engine/turn_heartbeat.rs:386-411— the heartbeat already records and reports(
:341detect,:426the existing "once per episode" contract); keep that, and make sure theescalation contract (report → turn ends) is what actually closes the turn, so a phase whose inner
bound failed (
:35-37) cannot outlive the report.interrupted/failed+ reason) and release the turn'sidle inhibitor, so a restart does not silently continue from a half-finished turn.
Key files
crates/tui/src/tui/ui/session_state.rs
crates/tui/src/tui/ui/dispatch.rs
crates/tui/src/core/engine/turn_heartbeat.rs
crates/tui/src/core/engine/turn_loop.rs
Acceptance criteria
still owns the turn" state.
interrupted/failed+ the stall reason) and releasesthe idle inhibitor.
a forced end.
Verification
cargo check -p codewhale-tui
cargo test -p codewhale-tui -- session_state
cargo test -p codewhale-tui -- dispatch
cargo test -p codewhale-tui -- turn_heartbeat
cargo clippy -p codewhale-tui -- -D warnings
New tests: recovery cancels the engine turn (assert the engine no longer owns a turn and the next
send is admitted); dispatch keeps the event loop responsive while a send is pending; an outcome is
persisted on recovery. The existing heartbeat contract test (
turn_heartbeat.rs:426) must keeppassing.
Out of scope
crates/tui/src/core/engine/turn_loop.rs:5737);this issue is about what happens after detection.
2026-09-29 13:59Z) and are already addressed in main by
c3e2c6fd8("keep undo under sizepressure…", 2026-09-29 16:25Z) — to be re-checked after a rebuild.
the async runtime), Expose stream retry budgets and transport timeouts as configuration #6700 (stream/transport budgets as configuration). The reporting half of this
arrived in 0.10.1 (
55b6dea79, "make a stalled turn report itself"); this is the acting half.Environment
operate; providercommand-code(deepseek/deepseek-v4.1-flash), OpenAI-compatible[tui] stream_max_resumes = 10,stream_max_transparent_retries = 10,stream_max_errors = 10,turn_wall_clock_secs = 86400(no global turn bound);stream_chunk_timeout_secsunset — theeffective inner bound in that session was 330 s (300 s + the 30 s
STALL_BOUND_GRACE)