Skip to content

feat(loop): execute marked native tool batches sequentially - #66

Open
Brian Krabach (bkrabach) wants to merge 3 commits into
mainfrom
feat/sequential-native-tool-batches-sonnet55
Open

Brian Krabach (bkrabach) wants to merge 3 commits into
mainfrom
feat/sequential-native-tool-batches-sonnet55

Conversation

@bkrabach

@bkrabach Brian Krabach (bkrabach) commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

What changed

  • Provider-marked _amplifier_execution_mode=sequential executes a native tool batch in provider response order.
  • A real vendor two-screenshot batch completed in order: start → finish, then start → finish.
  • When the first member fails, the second is skipped; both members' tool results are returned with is_error: true, including the computer-tagged result, and the provider continues to the next real request, which returned HTTP 200.
  • Added fresh internal goal evaluator/judge/summary calls with effort="high" and extended_thinking=false.
  • For those utility calls only, inherited budget/display kwargs are removed. Main-conversation role and model selection are unchanged.

Why

Native computer actions can share mutable UI state, so marked batches must run sequentially while ordinary unmarked batches retain their existing parallel behavior. Internal evaluator calls need compatible request settings when the global configuration uses xhigh; otherwise tool requests can inherit incompatible thinking settings. The utility-only override avoids that conflict without changing the main conversation.

Verification

  • Mac full suite: 486 passed.
  • Host-targeted tests: 93 passed.
  • Linux DTU joint gate: PASS, using the exact source at 8ce68dd8e2480f2804365ac9e31ef9c44e308c88 and the published, SHA-256-verified Core 2.0.1 wheel.
  • Live real-vendor two-screenshot batch verified ordered start → finish / start → finish execution.
  • Live first-failure probe verified second-member skip, paired error results through the provider, and continuation to the next real HTTP 200 request.

Qualification status

  • Qualified for the controlled synthetic/native sequential-batch scope described above, including the real-vendor two-screenshot round trip.
  • This does not claim qualification for all Mac write actions; those were not all tested.
  • The provider compatibility work in 1df remains separate in existing PR #130.
  • Downstream bundles were not released.
  • This remains a draft PR for CI and human review.

Breaking changes

No intended breaking changes. Sequential execution is opt-in through the provider marker; unmarked tool batches retain their parallel behavior. Utility-only evaluator settings do not change main-conversation role or model selection.

Optional managed-context handoff

  • Adds the optional paired context.final_request_record / context.signed_replay capability handoff. The loop snapshots the exact final request only after all overlays, budget fitting, measured-context transaction commits, and cancellation checks, and binds the admitted assistant response only after canonical context admission.
  • The pair is capability-gated: if either capability is absent, the existing behavior is unchanged. No Core kernel API change is required.
  • Overflow retries record only the request actually dispatched; rejected or cancelled requests are not admitted. The handoff covers normal foreground responses, streams, finalization, and provider overflow recovery.
  • This is an upstream loop capability only. Downstream Context/Unified/Live/managed changes remain separate and must not be published ahead of this merge.

Additional scoped verification

  • Reviewed source commit 4d8e4106c9f30e2f7dea53414ad8adb72cca345 is a direct child of the reviewed PR head 8ce68dd8e2480f2804365ac9e31ef9c44e308c88; no rebase, squash, or force-push was used.
  • Streaming-module qualification: 495 passed, 1 skipped, plus 10 reference integration/joint source-proof cases against published Core 2.0.1 on Python 3.13.15.
  • The managed-context consumer proof exercised both managed engines with the SDK-backed 5.5 path: genuine omitted-signature replay remained unchanged, the edited-prefix control was rejected as expected, portable summary options were accepted, canonical history was retained, and the post-compaction request succeeded. An explicit Sonnet 5 control also passed.
  • The optional handoff adds no regression for existing Core/context consumers; the existing PR behavior remains the default when the paired capabilities are not mounted.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Fresh internal goal judge, evaluator, and summary calls now request high reasoning effort and explicitly disable extended thinking without inheriting incompatible thinking budget or display settings. Main conversational requests remain unchanged.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach

Copy link
Copy Markdown
Collaborator Author

Review-readiness evidence (not a GitHub formal approval):

  • Source audit: PASS at 8ce68dd.
  • CI: all reported checks green on this exact head: Python 3.11, 3.12, 3.13, 3.14, and CLA.
  • Scoped release qualification: 3,427 passed, 21 skipped, 0 failures/errors against Core 2.0.1. Real native tool-batch execution was exercised, including ordered multi-member execution and fail-stop behavior.
  • Limitation: this is evidence for normal review; it is not a maintainer approval and makes no production-merged claim.

No CODEOWNERS or clear public prior reviewer assignment was found, and no reviewer is currently requested. Please use the repository's normal human approval gate.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants