Skip to content

feat: max_tokens becomes a real cap; max_tokens_fallback takes over the fallback - #42

Merged
Brian Krabach (bkrabach) merged 2 commits into
mainfrom
feat/max-tokens-cap
Sep 15, 2026
Merged

Brian Krabach (bkrabach) merged 2 commits into
mainfrom
feat/max-tokens-cap

Conversation

@bkrabach

Copy link
Copy Markdown
Collaborator

What changed

max_tokens was documented as "Maximum context size" but implemented as a priority-4 fallback in _calculate_budget — consulted only when a provider published no window. Orchestrators always pass a provider, so the knob was silently dead in production: the shipped foundation bundle's max_tokens: 300000 had no effect on when compaction fired, in either direction.

One overloaded key is split into the two jobs it was doing badly:

  • max_tokens (default None) is now a CAP, applied as min(derived_budget, max_tokens) on every derivation path with no exceptions. None means no cap — use the model's full window. A value above the model's own window is a no-op by construction, not an error.
  • max_tokens_fallback (default 200,000) takes over the fallback role. 200,000 is the smallest context window across the current generation of the three major vendors (Anthropic 200K base, OpenAI 272K default, Google ~1M); a guess that is too large overflows a request, one too small only compacts early.

Also fixes the usage log in add_message, which divided by max_tokens — a number that was never the governing denominator and is now usually unset. It now reports against the budget the last request actually ran on.

Why

The cadence probe had to patch this module's source in-container to add budget = min(budget, self.max_tokens) before either of its arms would compact at all. This ships that patch as the supported contract.

Breaking changes

None. Default None makes this a no-op for every existing config until someone opts in. An invalid cap (0, negative, non-int) degrades to "no cap" with a warning rather than being honored into a budget <= 0, which _should_compact's budget > 0 guard would turn into silently disabled compaction — matching how compaction_notice_token_reserve already handles a self-defeating value.

How to verify

test_compaction_trigger_provenance.py was inverted to pin the new contract in both directions — 45,000 caps low enough to compact where 70,000 does not, and an unset cap leaves the provider budget untouched. Plus validation coverage for the refused-cap path.

166 passed, 1 xfailed

Coordinated five-repo change

max_tokens in context-simple was documented as "Maximum context size" but implemented as a fallback consulted only when a provider published no context window. Orchestrators always pass a provider, so the knob was silently dead in production. It is now a real cap (default None = no cap); a new max_tokens_fallback (default 200,000) took over the fallback role.

Merge order matters

Merge FIRST (independent of each other, any order):

  1. amplifier-module-context-simple — feat/max-tokens-cap
  2. amplifier-module-provider-gemini — feat/publish-context-window
  3. amplifier-app-cli — feat/session-module-config-overrides

Merge AFTER those three:
4. amplifier-foundation — chore/drop-dead-max-tokens
5. amplifier-bundle-attractor — chore/drop-dead-max-tokens

Rationale: the two config repos delete max_tokens lines that only become safe-to-delete once context-simple's new semantics are in.

Cross-repo verification

DTU instance context-overflow-fix-20260915, all 8 target behaviors PASS:

  • A Gemini session's effective budget goes from 200,000 (the old fallback) to 1,011,712 (its real published window) — a 5.1x increase.
  • max_tokens: 500000 caps to exactly 500,000.
  • A cap above the model window is a no-op.
  • settings.yaml overrides.context-simple.config now reaches session.context through the CLI's own resolve_bundle_config.
Repo Result
context-simple 166 passed, 1 xfailed
provider-gemini 330 passed; 1 pre-existing failure (test_image_vision_integration_with_real_api, needs a live GOOGLE_API_KEY, fails on main too)
app-cli 2335 passed, 2 skipped, 1 xfailed
amplifier-foundation 1938 passed, 3 skipped; 1 pre-existing failure (test_grpc_adapter_main.py::TestVerifyModuleType::test_non_isinstance_object_with_mount_passes, verified failing on main before the change)
amplifier-bundle-attractor 268 passed, 2 skipped

…he fallback

`max_tokens` was documented as "Maximum context size" but implemented as
priority-4 fallback in `_calculate_budget` -- consulted only when a provider
published no window. Orchestrators always pass a provider, so the knob was
silently dead in production: the shipped foundation bundle's `max_tokens:
300000` had no effect on when compaction fired, in either direction. The
cadence probe had to patch this module's source in-container to add
`budget = min(budget, self.max_tokens)` before either of its arms would
compact at all.

This ships that patch as the supported contract, splitting one overloaded key
into the two jobs it was doing badly:

- `max_tokens` (default None) is now a CAP, applied as
  min(derived_budget, max_tokens) on every derivation path with no
  exceptions. None means no cap -- use the model's full window. A value above
  the model's own window is a no-op by construction, not an error.
- `max_tokens_fallback` (default 200,000) takes over the fallback role.
  200,000 is the smallest context window across the current generation of the
  three major vendors (Anthropic 200K base, OpenAI 272K default, Google ~1M);
  a guess that is too large overflows a request, one too small only compacts
  early.

Default None makes this a no-op for every existing config until someone opts
in. An invalid cap (0, negative, non-int) degrades to "no cap" with a warning
rather than being honored into a budget <= 0, which `_should_compact`'s
`budget > 0` guard would turn into silently disabled compaction -- matching
how `compaction_notice_token_reserve` already handles a self-defeating value.

Also fixes the usage log in `add_message`, which divided by `max_tokens` -- a
number that was never the governing denominator and is now usually unset. It
now reports against the budget the last request actually ran on.

Tests: `test_compaction_trigger_provenance.py` inverted to pin the new
contract in both directions -- 45,000 caps low enough to compact where 70,000
does not, and an unset cap leaves the provider budget untouched. Plus
validation coverage for the refused-cap path. 166 passed, 1 xfailed.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach

Copy link
Copy Markdown
Collaborator Author

The five PRs in this coordinated change

Merge FIRST — independent of each other, any order:

  1. context-simple — feat: max_tokens becomes a real cap; max_tokens_fallback takes over the fallback #42
  2. provider-gemini — fix: publish the real context window; refresh rates; purge the stale 8,192 amplifier-module-provider-gemini#48
  3. app-cli — fix: config overrides reach session.context and session.orchestrator amplifier-app-cli#342

Merge AFTER those three:
4. amplifier-foundation — microsoft/amplifier-foundation#388
5. amplifier-bundle-attractor — microsoft/amplifier-bundle-attractor#359

The two config repos (4, 5) delete max_tokens lines that only become safe-to-delete once context-simple's new cap semantics (1) are in.

… a log line

DTU validation of the max_tokens cap could not answer "what budget did this
session actually run on?" after the fact. _derive_budget logs the number it
picks, but the log never reaches the session record and did not surface on
stdout even at AMPLIFIER_LOG_LEVEL=INFO. That is the first question anyone
asks when compaction fires earlier than expected or a cap is suspected, and
it was unanswerable from the record.

Emits `context:budget` carrying the effective budget and its provenance:
source (provider_model_info | provider_defaults | max_tokens_fallback |
explicit), the window and reserve it was derived from, derived_budget,
max_tokens, max_tokens_fallback, capped, effective_budget.

Two properties matter as much as the number:

- Fires at the DELIVERY boundary, not at calculation time. A request that
  raises or is cancelled rolls back and must leave no observable trace --
  an invariant test_request_retention pins explicitly, and which the first
  version of this change broke: emitting on calculation made three rollback
  tests fail. Both delivery returns emit, including the common
  no-compaction path.
- Fires on CHANGE, not per request. The value is stable for most of a
  session, so per-request emission would drown the record it exists to make
  readable; a change (provider swap, mid-session cap, fallback kicking in) is
  always worth a line.

Observability is never load-bearing: no hooks, or an emitter that raises, is
logged and swallowed -- the same contract as context:compaction and
_emit_tool_result_ingress_truncation, both covered by tests here.

Verified against the real GeminiProvider: uncapped reports effective=1,011,712
source=provider_defaults capped=False; with max_tokens=500000 it reports
effective=500,000 capped=True.

Tests: 7 new in test_budget_event.py. test_request_retention's exact-event-list
assertion updated to include the new event, with a note on why exactly one
appears. 173 passed, 1 xfailed.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach

Copy link
Copy Markdown
Collaborator Author

New commit: context:budget — the effective budget becomes an artifact, not a log line

Pushed cd79245, which closes the one observability gap DTU validation of this PR surfaced.

The gap

DTU validation could not answer "what budget did this session actually run on?" after the fact. _derive_budget logs the number it picks, but that log never reaches the session record — and did not surface on stdout even at AMPLIFIER_LOG_LEVEL=INFO. That is the first question anyone asks when compaction fires earlier than expected, or when a cap is suspected, and it was unanswerable from the record.

This matters more after this PR than before it: now that max_tokens is a cap rather than the budget itself, the effective budget is derived from several inputs, so "which one won?" is a real question.

What it adds

A context:budget event carrying the effective budget and its provenance:

  • source — provider_model_info | provider_defaults | max_tokens_fallback | explicit
  • the window and reserve it was derived from
  • derived_budget, max_tokens, max_tokens_fallback, capped, effective_budget

Two design properties that matter as much as the number

  1. Fires at the delivery boundary, not at calculation time. A request that raises or is cancelled rolls back and must leave no observable trace — an invariant test_request_retention pins explicitly, and which the first version of this change broke: emitting on calculation made three rollback tests fail. Both delivery returns emit, including the common no-compaction path.
  2. Fires on change, not per request. The value is stable for most of a session, so per-request emission would drown the record it exists to make readable. A change — provider swap, mid-session cap, fallback kicking in — is always worth a line.

Observability is never load-bearing. No hooks, or an emitter that raises, is logged and swallowed — the same contract as context:compaction and _emit_tool_result_ingress_truncation, both covered by tests here.

Verification

Against the real GeminiProvider:

Config Reported
uncapped effective=1,011,712 source=provider_defaults capped=False
max_tokens=500000 effective=500,000 capped=True

Tests: 173 passed, 1 xfailed (7 new in test_budget_event.py). test_request_retention's exact-event-list assertion was updated to include the new event, with a note on why exactly one appears. ruff clean.

Merge order for this stack is unchanged — this PR still goes first.

Generated with Amplifier

@bkrabach
Brian Krabach (bkrabach) merged commit 85b4d08 into main Sep 15, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants