Fence stale supervisor writes and close stopped work on Continue - #2030
Merged
Merged
Conversation
ppXD
force-pushed
the
fix/fence-stale-supervisor-staging-on-continue
branch
from
September 30, 2026 00:39
4481d37 to
4ca7a08
Compare
A Continue can land while a stop is still in flight: the stopped walk, on another host, may still be deciding, claiming or staging its supervisor's next turn, and the stop's teardown runs only after the stop commits. The generation fence covered the engine's own parks, but the supervisor writes its decisions, waves and questions in its own DI scope, so a turn the Continue overtook still wrote under the revived run: a decision the revived walk replayed as its own, a wave it re-parked on, a question card a person could answer. The engine now carries the generation it claimed to every node it runs, and every supervisor write takes the park's share lock on the run at that generation inside its own transaction: the decision's claim, begin and terminal record, the spawn wave with its budget admission, so an overtaken turn reserves nothing, and an ask_human question, whose card, wait and decision record now commit as one, so a question anyone can see always has its token on the tape the ask API answers from. The terminal record takes the lock before it reads the decision's status: a walk whose decision the revived walk finished first stands down instead of throwing the illegal transition into the engine as the step failing. An overtaken turn stands down with RunSupersededException, which the engine does not record as a failure. A decision it claimed before the revive stays in flight, and the revived walk finishes it the way it finishes a crashed walk's. The revive also ends what the stopped attempt left, in its own transaction after the generation bump: it discards every pending wait and cancels the queued and running agents and the staged child runs. The running agents' kills and the revived run's dispatch run once it commits, on no request token, so a client that goes away after the commit cannot cut them short. Before, the revived supervisor re-parked on the stopped wave, reclaimed its queued agent or folded a still-running agent onto its tape as "Running" for good, and a map branch adopted its old wait. The fold now also leaves a wave alone while any of its agents is live. A replayed spawn decision whose wave the stop closed restages only the closed slots: a finished agent keeps its answered wait, and each slot keeps the budget reservation it was first admitted under, which a recomputed deadline no longer turns into a refusal. The stop's teardown now skips waits another transaction holds, passes over them again while that transaction may still let them go, and runs each step on its own, so a Continue holding those waits can no longer deadlock it out of its kill-wave, nor leave one open by rolling back.
ppXD
force-pushed
the
fix/fence-stale-supervisor-staging-on-continue
branch
from
September 30, 2026 13:53
4ca7a08 to
ea3b2dc
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A supervisor turn a Continue overtook writes nothing under the revived run. The stopped walk (on another host, so the stop never tripped it) can still be deciding, claiming or staging its next turn when the Continue lands. The engine carries the generation it claimed to every node it runs (
RunGenerationFence.ClaiminWorkflowEngine.RunNodeOnceAsync). Every supervisor write takes the engine park'sSELECT … FOR SHAREon the run at that generation inside its own transaction (RunGenerationFence.EnterAsync/CommitUnderClaimAsync):SupervisorDecisionLog);RealSupervisorActionExecutor.StageAgentsAndParkAsync);SupervisorTurnService.ExecuteAndRecordTogetherAsync). A question anyone can see therefore always has its token on the tape thatISupervisorAskAnswerService, the plan confirmation and the human-touch reader answer from.The terminal record takes the fence before it reads the decision's status. A walk whose decision the revived walk finished first stands down, instead of throwing an illegal Succeeded → Succeeded into the engine's generic catch as the supervisor step failing. An overtaken turn throws
RunSupersededException, which the engine lets through as the walk's end, never as a node failure. A decision it claimed before the revive stays in flight, and the revived walk finishes it once, as it finishes a crashed walk's.Supersededdecision status. It needs a migration and a partial unique index, because the revived walk's deterministic re-decision collides on the (run, key) index, and the Duplicate path would replay the superseded row as an empty turn that is never counted. Also rejected for the ask: skipping the node's human re-entry guard while an ask is in flight, so a replay records the token late. That leaves a window where a posted question has no token on the tape; the single transaction has none.Budget admission runs inside the fenced wave transaction (
AdmitWaveAsync, afterEnterAsync; the ledger joins the transaction). An overtaken turn reserves nothing, and a refusal rolls back with the wave. A slot that already holds a row (a closed slot restaged, or a finished slot kept) keeps the reservation it was first admitted under. Before, a re-reservation computed a new deadline, whichReserveAsyncreads as a different intent, so a capped run's restage was refused as over budget. Rejected: tolerating a recomputed deadline, or re-reserving a Released row, in the ledger; both break its pinned identical-intent contract.Continue ends what the stopped attempt left, in the revive's transaction after the bump (
CloseEndedAttemptAsync). It discards every Pending wait, cancels Queued agents, cancels staged Pending/Enqueued child runs, and CASes Running agents to Cancelled with the row half of the cancel (IRunningAgentCancellation.CancelRunningRowAsync).FinishRunningCancelAsync, one post-commit action per agent). They run outside any transaction, as the decision lock order from Close an ended agent run's decisions and cancel started sub-workflows #2029 requires. They also run onCancellationToken.None, as does the revived run's dispatch, so a client that disconnects after the commit cannot cut them short.IAgentRunService.A closed wave restages only its closed slots. A finished agent keeps its answered wait and runs no second time (
TurnWave;DropClosedWaveAsyncdeletes Discarded rows only).The stop's teardown can no longer be deadlocked out of its kill-wave, or leave a wait open. A Continue's revive holds every Pending wait of the run until it commits, and it and the teardown locked those rows in opposite orders.
FOR UPDATE SKIP LOCKED), then pass over them again, up to three times 200 ms apart (CloseWaitsSkippingHeldOnesAsync). A holder that rolls back instead of closing them therefore releases them to a later pass, rather than leaving them Pending for the kill to answer. What stays held is logged.TearDownStepAsync), so a failed step never costs the steps after it.Accepted residuals:
ChooseDecisionAsync, the rehydrate fold's outcome write andReopenDiscardedAskAsyncare not fenced.Follow-ups:
WorkflowEngine.CancelPendingWaitsAndChildrenAsync) is still unfenced on the generation, so a slow ceremony can discard waits the revived walk already parked.AgentReviewRunner), created after the revive, runs under the live revived run, and nothing kills it.Test plan
New and extended integration tests over real Postgres. Each drives either the real engine walk on another host's cancellation registry, or the real turn service under that walk's claim, with the hold point named in each test. A fixture-root decoration of the decision log (
ScriptedDecisionLog, armed per run bySupervisorDecisionScript.HoldDecisionLog) lets a test hold the engine's own walk inside the supervisor node.SupervisorStopContinueFlowTests(the wave cases run both capped and uncapped):A_turn_a_continue_overtook_while_it_decided_……claims_nothing_after_the_revived_turn_decided_otherwiseA_stop_teardown_landing_after_a_continue_finds_nothing_…A_stopped_wave_whose_decision_was_still_in_flight_…A_stopped_wave_staged_afresh_keeps_the_agent_that_had_already_finishedA_turn_that_claimed_its_decision_before_the_stop_begins_nothing_after_the_continueA_turn_already_assembling_its_wave_when_the_continue_lands_reserves_nothingA_turn_whose_wave_committed_before_the_stop_records_no_terminal_after_the_continueA_stop_landing_while_a_wave_is_staged_waits_for_the_whole_wave_…A_walk_whose_decision_the_revived_walk_finished_first_stands_down_without_failing_the_step: through the engine, with the revived walk recording the plan before the overtaken walk's executor returnsA_continue_whose_request_went_away_after_its_commit_still_ends_the_running_agents_and_dispatches_the_runSupervisorAskHumanStopContinueFlowTests, each answering throughISupervisorAskAnswerService, never the wait's token:An_ask_a_continue_overtook_before_its_question_posts_no_card_parks_no_wait_and_records_nothingA_continue_landing_while_an_ask_posts_its_card_waits_for_it_and_the_revived_run_answers_from_that_cardA_gate_question_asked_with_no_conversation_as_a_continue_lands_stays_answerable_through_the_ask_apiContinueParkedRunFlowTests:A_continue_holding_the_stopped_attempts_waits_never_stalls_the_stop_teardown_short_of_its_kill_waveA_teardown_step_that_fails_never_costs_the_kill_wave(a 40P01 injected into the close)A_wait_the_teardown_had_to_skip_is_closed_once_the_transaction_holding_it_lets_go(a second connection holds the wait, then lets go)OperatorCancelInProgressWalkFlowTests: the park lock-hold test now asserts the stopped walk's outcome and releases its hold in afinally. It uses the kit's helpers, whose lock-wait check fails at once when the writer finishes without waiting.Unit:
RunGenerationFenceTests(claim read back for its run only, nested restore, async flow, lock refused outside a transaction, no-op outside a walk) and five fold-guard cases inSupervisorAgentResultsFoldTests.Mutations, each red on this head:
…begins_nothing_after_the_continue…records_no_terminal_after_the_continue,A_walk_whose_decision_the_revived_walk_finished_first…A_walk_whose_decision_the_revived_walk_finished_first…(anode.failedlands in the revived run)…reserves_nothing(capped)A_stop_teardown_landing_after_a_continue_finds_nothing_…,…request_went_away……request_went_away……never_stalls_the_stop_teardown…,…skip_is_closed_once……never_costs_the_kill_wave…skip_is_closed_once……keeps_the_agent_that_had_already_finishedA_stop_landing_while_a_wave_is_staged…RunGenerationFenceTestsM1d (no stand-down read before the re-park) survives by design: the fenced terminal record already refuses what the re-park does next.
SupervisorStopContinueFlowTests19/19,SupervisorAskHumanStopContinueFlowTests3/3,ContinueParkedRunFlowTests10/10,OperatorCancelInProgressWalkFlowTests15/15Every non-real-model
Workflows.Supervisor*integration class 694/694; budget ledger and settlement classes 94/94; cancel, continue, decision, sub-workflow and reconciler neighbours 634/634Unit suite: 11630 passed, 1 skipped
dotnet build CodeSpace.sln: 0 errors