fix: a stored codex failure is not adoptable by a later dispatch - #402
Conversation
Two intake codex runs timed out against an unavailable upstream. Their failures were written to the codex-adoption records, and every later dispatch for the same prompt hash adopted the stored failure instead of running codex. `workflow_select` then dead-lettered continuously, reporting a 3600s timeout that took five seconds, and no retry could ever clear it. The engine already treats every codex failure as environmental: `boundary_resource::class_for_adapter_failure` has no deterministic-failure class and defaults to `provider-unavailable`. A stored failure therefore describes the environment at the time of that run, not a property of the prompt, and must not answer a dispatch that did not join the run producing it. - `run_adoptable_codex_request` reuses a completed record only when the run succeeded. `wait_for_adoption_result` is unchanged: a waiter attached to a run in flight is entitled to that run's failure, and rejecting it there would make waiters hang. - The outcome predicate is the exit code, not `error_kind`. `read_completed_adoption_result` builds records without `error_kind` through `CodexResult::success` even when the process exited non-zero. - `start_adoption_worker` discards the previous run's `result.json`, `effect.json`, and `effect-receipts.log` before publishing its intent. Those three are the recovery paths; without this a waiter on the fresh run promotes the previous failure straight back and the fix does nothing. `spawn_codex_sync_preserves_typed_jsonl_refusal_through_adoption` asserted that two dispatches spawn codex once. Its fixture emits `turn.failed` / "server overloaded" — the environmental shape this change is about — so that incidental assertion encoded the reported defect. The test's actual subject, typed metadata surviving the adoption machinery, is unchanged. Refs ChronoAIProject/fkst-packages#3919 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The devloop refused ChronoAIProject/fkst-packages#3919 as `wrong-layer` and named the gap in the first commit on this branch: it replaces completed failures, but `CodexResult`/`into_lua_table` still expose neither `run_id` nor produced/adopted provenance, so no contract-complete substrate revision exists for packages to pin. #3919 asks for both halves. This is the second: an adopted five-second read and a real hour-long run reported the same thing, which is why identifying the defect took a measurement of `ELAPSED_MS` rather than a reading of the message. - `CodexResult` carries `provenance` (`produced` / `adopted`) and the `run_id` of the run that produced the record; `into_lua_table` exposes both. - Provenance is decided by comparing the stored `run_id` against the requesting run's, not by which code path returned the result. In the adoption protocol the producing call also reads its own outcome back through `read_completed_adoption_result`, so a path-based rule marks a producer `adopted`. The added test caught exactly that. - The requesting run id is threaded through the read and recover paths. Refs ChronoAIProject/fkst-packages#3919 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
Second commit pushed:
That refusal is correct, and #3919 does ask for both halves. The first commit did the replay half only. What the second commit adds
The path-based rule was wrong and a test caught itMy first attempt marked provenance by code path — anything returned from Verification
⟦AI:FKST⟧ |
Root-cause fix for ChronoAIProject/fkst-packages#3919.
What was happening
Two intake codex runs timed out against an unavailable upstream. Their failures were written to the codex-adoption records, and every later dispatch for the same prompt hash adopted the stored failure instead of running codex.
workflow_selectdead-lettered continuously — measured at 57–62 failures per 30 minutes — reportingcodex timed out after 3600s wall clockon runs that took four seconds. No retry could clear it, because the path that would produce a fresh result was never taken.Why failures must not be adoptable
The engine already classifies every codex failure as environmental.
boundary_resource::class_for_adapter_failurehas no deterministic-failure class; its default arm isprovider-unavailable. A stored failure describes the environment at the time of that run, not a property of the prompt, so it must not answer a dispatch that did not join the run producing it. This follows the existing SPEC contract rather than adding a new expiry mechanism.The change
run_adoptable_codex_requestreuses a completed record only when the run succeeded.wait_for_adoption_resultis deliberately unchanged: a waiter attached to a run in flight is entitled to that run's failure, and rejecting it there would make waiters hang until their own deadline.error_kind.read_completed_adoption_resultbuilds records withouterror_kindthroughCodexResult::successeven when the process exited non-zero — the first version of this patch usederror_kindand the new test caught it.start_adoption_workerdiscards the previous run'sresult.json,effect.json, andeffect-receipts.logbefore publishing its intent. Those are the three recovery paths inrecover_completed_adoption_result_locked_from_disk; without this step a waiter on the fresh run promotes the previous failure straight back and the fix does nothing.An existing assertion changed — please look at this one
spawn_codex_sync_preserves_typed_jsonl_refusal_through_adoptionasserted that two dispatches spawn codex exactly once. Its fixture emits{"type":"turn.failed","error":{"message":"server overloaded",…}}— the environmental shape this change is about — so that assertion encoded the defect #3919 reports. It is now"2", with the reason in a comment. The test's actual subject, typed metadata surviving the adoption machinery, is untouched and still asserted on every run.Verification
cargo test -p fkst-framework --test sdk_codex— 167 passed, 0 failed.scripts/run.sh test—FKST_LOCAL_ITERATION_RESULT:v2:PASS:NONE, twice (before and after the SPEC sentence).left: 1, right: 0.This PR cannot pass the devloop review gate right now
llm.aelf.devPOST /responseshas returned 503 for every one of 120 consecutive probes (sub-second refusal, all ten text models,GET /modelsreturns 200). Every review-meta judgment times out, which is why #400 has been stalled infixingsince 2026-08-18T04:34Z. This PR will stall in the same place until the upstream recovers. It is opened for human review, not for the devloop to merge.⟦AI:FKST⟧
🤖 Generated with Claude Code