Skip to content

fix: a stored codex failure is not adoptable by a later dispatch - #402

Merged
macstudio-4[bot] merged 2 commits into
integration-macstudio-4from
fix/adoption-failure-ttl
Aug 21, 2026
Merged

macstudio-4[bot] merged 2 commits into
integration-macstudio-4from
fix/adoption-failure-ttl

Conversation

@macstudio-4

@macstudio-4 macstudio-4 Bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

Root-cause fix for ChronoAIProject/fkst-packages#3919.

What was happening

Two intake codex runs timed out against an unavailable upstream. Their failures were written to the codex-adoption records, and every later dispatch for the same prompt hash adopted the stored failure instead of running codex. workflow_select dead-lettered continuously — measured at 57–62 failures per 30 minutes — reporting codex timed out after 3600s wall clock on runs that took four seconds. No retry could clear it, because the path that would produce a fresh result was never taken.

Why failures must not be adoptable

The engine already classifies every codex failure as environmental. boundary_resource::class_for_adapter_failure has no deterministic-failure class; its default arm is provider-unavailable. A stored failure describes the environment at the time of that run, not a property of the prompt, so it must not answer a dispatch that did not join the run producing it. This follows the existing SPEC contract rather than adding a new expiry mechanism.

The change

  • run_adoptable_codex_request reuses a completed record only when the run succeeded. wait_for_adoption_result is deliberately unchanged: a waiter attached to a run in flight is entitled to that run's failure, and rejecting it there would make waiters hang until their own deadline.
  • The outcome predicate is the exit code, not error_kind. read_completed_adoption_result builds records without error_kind through CodexResult::success even when the process exited non-zero — the first version of this patch used error_kind and the new test caught it.
  • start_adoption_worker discards the previous run's result.json, effect.json, and effect-receipts.log before publishing its intent. Those are the three recovery paths in recover_completed_adoption_result_locked_from_disk; without this step a waiter on the fresh run promotes the previous failure straight back and the fix does nothing.

An existing assertion changed — please look at this one

spawn_codex_sync_preserves_typed_jsonl_refusal_through_adoption asserted that two dispatches spawn codex exactly once. Its fixture emits {"type":"turn.failed","error":{"message":"server overloaded",…}} — the environmental shape this change is about — so that assertion encoded the defect #3919 reports. It is now "2", with the reason in a comment. The test's actual subject, typed metadata surviving the adoption machinery, is untouched and still asserted on every run.

Verification

  • cargo test -p fkst-framework --test sdk_codex — 167 passed, 0 failed.
  • scripts/run.sh test — FKST_LOCAL_ITERATION_RESULT:v2:PASS:NONE, twice (before and after the SPEC sentence).
  • The new test is a positive control: it fails on the pre-change engine with left: 1, right: 0.

This PR cannot pass the devloop review gate right now

llm.aelf.dev POST /responses has returned 503 for every one of 120 consecutive probes (sub-second refusal, all ten text models, GET /models returns 200). Every review-meta judgment times out, which is why #400 has been stalled in fixing since 2026-08-18T04:34Z. This PR will stall in the same place until the upstream recovers. It is opened for human review, not for the devloop to merge.

⟦AI:FKST⟧

🤖 Generated with Claude Code

macstudio-4[bot] and others added 2 commits August 20, 2026 04:59
Two intake codex runs timed out against an unavailable upstream. Their
failures were written to the codex-adoption records, and every later
dispatch for the same prompt hash adopted the stored failure instead of
running codex. `workflow_select` then dead-lettered continuously, reporting
a 3600s timeout that took five seconds, and no retry could ever clear it.

The engine already treats every codex failure as environmental:
`boundary_resource::class_for_adapter_failure` has no deterministic-failure
class and defaults to `provider-unavailable`. A stored failure therefore
describes the environment at the time of that run, not a property of the
prompt, and must not answer a dispatch that did not join the run producing it.

- `run_adoptable_codex_request` reuses a completed record only when the run
  succeeded. `wait_for_adoption_result` is unchanged: a waiter attached to a
  run in flight is entitled to that run's failure, and rejecting it there
  would make waiters hang.
- The outcome predicate is the exit code, not `error_kind`.
  `read_completed_adoption_result` builds records without `error_kind`
  through `CodexResult::success` even when the process exited non-zero.
- `start_adoption_worker` discards the previous run's `result.json`,
  `effect.json`, and `effect-receipts.log` before publishing its intent.
  Those three are the recovery paths; without this a waiter on the fresh run
  promotes the previous failure straight back and the fix does nothing.

`spawn_codex_sync_preserves_typed_jsonl_refusal_through_adoption` asserted
that two dispatches spawn codex once. Its fixture emits `turn.failed` /
"server overloaded" — the environmental shape this change is about — so that
incidental assertion encoded the reported defect. The test's actual subject,
typed metadata surviving the adoption machinery, is unchanged.

Refs ChronoAIProject/fkst-packages#3919

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The devloop refused ChronoAIProject/fkst-packages#3919 as `wrong-layer` and
named the gap in the first commit on this branch: it replaces completed
failures, but `CodexResult`/`into_lua_table` still expose neither `run_id` nor
produced/adopted provenance, so no contract-complete substrate revision exists
for packages to pin.

#3919 asks for both halves. This is the second: an adopted five-second read and
a real hour-long run reported the same thing, which is why identifying the
defect took a measurement of `ELAPSED_MS` rather than a reading of the message.

- `CodexResult` carries `provenance` (`produced` / `adopted`) and the `run_id`
  of the run that produced the record; `into_lua_table` exposes both.
- Provenance is decided by comparing the stored `run_id` against the requesting
  run's, not by which code path returned the result. In the adoption protocol
  the producing call also reads its own outcome back through
  `read_completed_adoption_result`, so a path-based rule marks a producer
  `adopted`. The added test caught exactly that.
- The requesting run id is threaded through the read and recover paths.

Refs ChronoAIProject/fkst-packages#3919

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@macstudio-4

macstudio-4 Bot commented Aug 20, 2026

Copy link
Copy Markdown
Contributor Author

Second commit pushed: 5e7498c. It closes the gap the devloop named when it refused ChronoAIProject/fkst-packages#3919 as wrong-layer:

That candidate replaces completed failures but CodexResult/into_lua_table still expose neither run_id nor produced/adopted provenance. Packages owns only the pin, so no contract-complete accepted substrate revision exists to pin.

That refusal is correct, and #3919 does ask for both halves. The first commit did the replay half only.

What the second commit adds

  • CodexResult carries provenance (produced / adopted) and the run_id of the run that produced the record. into_lua_table exposes both.
  • Provenance is decided by comparing the stored run_id against the requesting run's id, not by which code path returned the result. The requesting run id is threaded through the read and recover paths.

The path-based rule was wrong and a test caught it

My first attempt marked provenance by code path — anything returned from read_completed_adoption_result was adopted. That marks the producer adopted too, because in the adoption protocol the call that spawned the worker also reads its own outcome back through that same function. The added test failed on its first assertion (left: "adopted", right: "produced") and forced the run-id comparison. Reading the code did not surface it.

Verification

  • cargo test -p fkst-framework --test sdk_codex — 168 passed, 0 failed.
  • scripts/run.sh test — FKST_LOCAL_ITERATION_RESULT:v2:PASS:NONE.
  • SPEC.md states the contract for both halves.

⟦AI:FKST⟧

@macstudio-4
macstudio-4 Bot merged commit ec3c4c1 into integration-macstudio-4 Aug 21, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants