Skip to content

fix(loop): recover provider input overflow once. - #56

Merged
Brian Krabach (bkrabach) merged 2 commits into
mainfrom
fix/provider-overflow-recovery
Sep 15, 2026
Merged

Brian Krabach (bkrabach) merged 2 commits into
mainfrom
fix/provider-overflow-recovery

Conversation

@bkrabach

Copy link
Copy Markdown
Collaborator

What changed

Adds one bounded recovery path for a provider-reported input-context overflow:

  • An optional synchronous recover_context_overflow capability can return a strict budget dictionary after a ContextLengthError.
  • The loop accepts only non-boolean integer budget fields, observed input above the provider allowance, and a positive context target below the failed estimate.
  • It performs at most one hard-fit rebuild, one retry preflight, and one retry before any outer streaming chunk has been yielded. A None retry-preflight result remains unproven rather than being treated as a local fit; the prior server rejection is the narrow authorization for that retry.
  • Resolved retention and request overlays are replayed without rerunning hooks or tools. Retries retain the same or lower output cap; finalization retains tool_choice="none"; cancellation, iterator finalization, and tool-turn closure remain protected.
  • request_options are forwarded only to optional methods that explicitly declare that named parameter. Providers with only **kwargs retain their existing call shape.
  • A scalar overflow-recovery observability event distinguishes unavailable feedback from invalid feedback.

Why

A provider can reject an input after ordinary request assembly. The generic loop needs a narrowly bounded, provider-authorized way to retry with a smaller context without parsing provider-private error text or replaying turn side effects.

Compatibility and scope

Providers that do not expose the recovery capability retain the existing dispatch behavior. The generic loop does not define or parse provider-private feedback formats. This change introduces no release or host-rollout claim.

Validation

Previously completed isolated validation against Core 1.6.1 (not rerun while publishing this PR):

  • Loop: 340 passed, with 1 optional cross-module collection skip.
  • Existing OpenAI + Context plus new Loop joint collection: 2 passed.
  • Combined Anthropic collection: 848 passed, with 7 live/key-dependent skips.
  • Total: 1,190 passed. Nine paired SDK-fake cases also passed their respective assertions.
  • The same harness exercised cold API overflow (one rebuild, one retry, then four ordinary requests), warm pre-SDK overflow (one hard-fit plus four ordinary requests without resurrection), and negative budget bounds.

SDK/transport cases use a synthetic byte-count oracle, not native vendor token limits; they do not prove live restoration or general output quality. Remaining module-wide Ruff diagnostics are pre-existing, so this is not a full clean-lint claim.

Follow-up

A companion Anthropic change remains pending Loop main CI; it is not part of this PR.

How to verify

Run the module's provider budget and overflow recovery tests in an environment with the compatible Core version, then run the applicable joint provider/context collection.

Breaking changes

None expected for providers without the optional recovery capability.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants