Skip to content

chore: drop the dead max_tokens from every bundle; restore minimal-delegate's 64k - #388

Merged
Brian Krabach (bkrabach) merged 2 commits into
mainfrom
chore/drop-dead-max-tokens
Sep 15, 2026
Merged

Brian Krabach (bkrabach) merged 2 commits into
mainfrom
chore/drop-dead-max-tokens

Conversation

@bkrabach

Copy link
Copy Markdown
Collaborator

Merge this AFTER context-simple, provider-gemini and app-cli. These deletions are only safe once context-simple's new semantics are in.

What changed — two commits

1. chore: drop the dead max_tokens from every bundle

context-simple's max_tokens used to be a fallback consulted only when a provider published no context window. Orchestrators always pass a provider, so these values never reached the wire — they were inert in every bundle here.

It is now a CAP defaulting to None (no cap). Leaving these lines in place would silently clamp every session to 300,000 tokens — including sessions on 1M-context models — turning a documentation fix into a 3x behavior change nobody asked for. Deleting them preserves today's behavior exactly, and leaves no fossil that looks meaningful.

Files, all session.context.config.max_tokens:

File Value removed
bundle.md 300000
bundles/anchors/bundle.md 300000
experiments/delegation-only/bundle.md 300000
experiments/exp-foundation.md 300000
experiments/exp-lean/behaviors/lean-foundation.yaml 300000
experiments/minimal-delegate/.../minimal-delegate-foundation.yaml 64000

anchors-amp-dev, amplifier-dev, minimal, with-anthropic and with-openai set no max_tokens of their own — they inherit, so the root bundle.md covers them.

2. feat(minimal-delegate): restore the 64k budget, now that it actually applies

That experiment asked for max_tokens: 64000 — a deliberately tight budget, the whole point of a delegate-only head — and never got it. The previous commit deleted the line to preserve behavior; restoring it is a deliberate choice, not an inherited one: the cap works, nobody depends on this bundle today, and the 64k constraint is finally measurable.

How to verify

Every touched file was re-parsed after the edit to prove the YAML still loads and only the intended key is gone.

1938 passed, 3 skipped

1 pre-existing failure, verified failing on main before this change: test_grpc_adapter_main.py::TestVerifyModuleType::test_non_isinstance_object_with_mount_passes.

Breaking changes

Behavior is preserved exactly for every bundle except minimal-delegate, which now genuinely runs under a 64,000-token cap — an intentional, opted-in change to an experiment with no current dependents.

Coordinated five-repo change

max_tokens in context-simple was documented as "Maximum context size" but implemented as a fallback consulted only when a provider published no context window. Orchestrators always pass a provider, so the knob was silently dead in production. It is now a real cap (default None = no cap); a new max_tokens_fallback (default 200,000) took over the fallback role.

Merge order matters

Merge FIRST (independent of each other, any order):

  1. amplifier-module-context-simple — feat/max-tokens-cap
  2. amplifier-module-provider-gemini — feat/publish-context-window
  3. amplifier-app-cli — feat/session-module-config-overrides

Merge AFTER those three:
4. amplifier-foundation — chore/drop-dead-max-tokens
5. amplifier-bundle-attractor — chore/drop-dead-max-tokens

Rationale: the two config repos delete max_tokens lines that only become safe-to-delete once context-simple's new semantics are in.

Cross-repo verification

DTU instance context-overflow-fix-20260915, all 8 target behaviors PASS:

  • A Gemini session's effective budget goes from 200,000 (the old fallback) to 1,011,712 (its real published window) — a 5.1x increase.
  • max_tokens: 500000 caps to exactly 500,000.
  • A cap above the model window is a no-op.
  • settings.yaml overrides.context-simple.config now reaches session.context through the CLI's own resolve_bundle_config.
Repo Result
context-simple 166 passed, 1 xfailed
provider-gemini 330 passed; 1 pre-existing failure (test_image_vision_integration_with_real_api, needs a live GOOGLE_API_KEY, fails on main too)
app-cli 2335 passed, 2 skipped, 1 xfailed
amplifier-foundation 1938 passed, 3 skipped; 1 pre-existing failure (test_grpc_adapter_main.py::TestVerifyModuleType::test_non_isinstance_object_with_mount_passes, verified failing on main before the change)
amplifier-bundle-attractor 268 passed, 2 skipped

context-simple's `max_tokens` used to be a fallback consulted only when a
provider published no context window. Orchestrators always pass a provider, so
these values never reached the wire -- they were inert in every bundle here.

It is now a CAP, defaulting to None (no cap). Leaving these lines in place
would silently clamp every session to 300,000 tokens -- including sessions on
1M-context models -- turning a documentation fix into a 3x behavior change
nobody asked for. Deleting them preserves today's behavior exactly, and leaves
no fossil that looks meaningful.

Files, all `session.context.config.max_tokens`:
  bundle.md                                                 300000
  bundles/anchors/bundle.md                                 300000
  experiments/delegation-only/bundle.md                     300000
  experiments/exp-foundation.md                             300000
  experiments/exp-lean/behaviors/lean-foundation.yaml       300000
  experiments/minimal-delegate/.../minimal-delegate-foundation.yaml 64000

anchors-amp-dev, amplifier-dev, minimal, with-anthropic and with-openai set no
max_tokens of their own -- they inherit, so the root bundle.md covers them.

WORTH A DECISION, deliberately not made here: minimal-delegate asked for
64,000, a tight budget matching the experiment's intent, and never got it.
That intent is now expressible for the first time. Restoring it is a
behavior change on an experiment, so it should be chosen, not inherited from
a line that never worked.

Every touched file re-parsed after the edit to prove the YAML still loads and
only the intended key is gone.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
…applies

This experiment asked for `max_tokens: 64000` -- a deliberately tight budget,
the whole point of a delegate-only head -- and never got it: `max_tokens` was
a fallback consulted only when a provider published no window, and the
orchestrator always passes one.

The previous commit deleted the line to preserve behavior. Restoring it now is
a deliberate choice, not an inherited one: the cap works, nobody depends on
this bundle today, and the 64k constraint is finally measurable.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach

Copy link
Copy Markdown
Collaborator Author

The five PRs in this coordinated change

Merge FIRST — independent of each other, any order:

  1. context-simple — feat: max_tokens becomes a real cap; max_tokens_fallback takes over the fallback amplifier-module-context-simple#42
  2. provider-gemini — fix: publish the real context window; refresh rates; purge the stale 8,192 amplifier-module-provider-gemini#48
  3. app-cli — fix: config overrides reach session.context and session.orchestrator amplifier-app-cli#342

Merge AFTER those three:
4. amplifier-foundation — #388
5. amplifier-bundle-attractor — microsoft/amplifier-bundle-attractor#359

The two config repos (4, 5) delete max_tokens lines that only become safe-to-delete once context-simple's new cap semantics (1) are in.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants