Skip to content

fix: checkpoint settled tool batches before dispatch - #60

Merged
Brian Krabach (bkrabach) merged 1 commit into
mainfrom
fix/ykq-durable-checkpoint-loop
Sep 18, 2026
Merged

Brian Krabach (bkrabach) merged 1 commit into
mainfrom
fix/ykq-durable-checkpoint-loop

Conversation

@bkrabach

Copy link
Copy Markdown
Collaborator

What changed

  • Adds the optional zero-argument session.durable_checkpoint coordinator capability.
  • Invokes that capability after a normal concurrent tool batch has settled and its results have been appended in original order, before another provider request or bounded-budget finalization.
  • Supports synchronous and awaitable callbacks. A non-callable registration, literal False, or ordinary exception stops further provider dispatch with an explicit error; cancellation propagates.
  • Documents the capability and adds focused coverage for ordering, finalization, failure, cancellation, no-capability, and no-tool behavior.

Why

A later provider request can fail after tool work completed. Hosts that opt in can now durably persist the completed, canonical batch at the boundary where it is safe to resume.

Compatibility and scope

  • Hosts without the optional capability keep their existing behavior.
  • The Loop does not write transcripts; the host owns persistence.
  • End-to-end CLI persistence requires a compatible CLI host to register this capability. This PR is the upstream Loop half and does not, by itself, claim that end-to-end CLI behavior.
  • No Core, Context, provider, or cancellation-policy changes.

Verification

  • pytest tests -q — 389 passed.
  • Focused coverage exercises one- and two-tool ordered batches, normal continuation, bounded-budget finalization, callback failures, cancellation propagation, absent capability, and no-tool turns.
  • The compatible CLI host change was validated with the paired Loop change locally; this Loop-only PR does not assert that the module alone provides CLI transcript persistence.

Breaking changes

None for existing hosts. Hosts that register session.durable_checkpoint now receive an explicit failure instead of a further provider dispatch when their checkpoint cannot complete, by design.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach

Copy link
Copy Markdown
Collaborator Author

Independent Fable review — MERGE-READY

No correctness or failure-policy blockers found. This is a review-summary record, not a formal approval.

Reviewed Loop PR #60 at 939f5d1 and paired local, unpushed CLI commit 5c0beb3; no CLI PR exists, so this is not a CLI PR review. Model execution confirmed 32/32 successful direct responses from claude-fable-5-1; instance/account identity was not established.

The review supports one settled, ordered, post-append checkpoint before normal continuation and budget finalization; an always-writing CLI handler; failure stopping redispatch; cancellation propagation; and compatibility for hosts without the optional capability.

Nonblocking findings — not changes in this PR:

  1. Include underlying disk-error detail in the outer nonverbose CLI checkpoint error (the cause is already chained).
  2. Consider defensive get_capability lookup for non-Core host doubles; the real ModuleCoordinator has it and CI passes.
  3. Checkpoint and final-save metadata model labels differ cosmetically/inheritedly.
  4. resumable=False error metadata/comment predates the new partially persisted turn semantics; no behavior change proposed.
  5. The paired CLI README could explicitly say “after a completed tool batch.”

No runtime changes are requested as merge prerequisites; keep the verified code unchanged.

Verification: exact HEAD blobs match the tested R5 candidate; Loop 389 passed; paired CLI 2,362 passed / 3 skipped / 1 expected failure; CLI integration 13 passed; focused lint/format passed. Six synthetic overflow and three healthy real CLI/Core/Loop/Context/SessionStore cases preserved exact IDs, order, and values on disk and fresh resume; three typed PTY checks found disk identical before and after exit. Published Loop CI passed Python 3.11–3.14 and CLA.

This is not proof against remote API/token-window/protected-floor behavior or normal unwrapped-bundle configuration; combined runtime used Linux/Python 3.12/Core 1.6.1 while CI uses Core 1.5.2. Interactive disk-write-failure recovery was not exercised end-to-end (unit failure-policy coverage only).

Release boundary: upstream Loop half only; CLI remains local pending Loop merge and green main CI. No merge, deployment, or authority change is performed by this comment.

@bkrabach
Brian Krabach (bkrabach) merged commit 603aa6e into main Sep 18, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants