Skip to content

chore: drop the dead max_tokens from both e2e pipeline profiles - #359

Merged
Brian Krabach (bkrabach) merged 1 commit into
mainfrom
chore/drop-dead-max-tokens
Sep 15, 2026
Merged

Brian Krabach (bkrabach) merged 1 commit into
mainfrom
chore/drop-dead-max-tokens

Conversation

@bkrabach

@bkrabach Brian Krabach (bkrabach) commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Merge this AFTER context-simple, provider-gemini and app-cli. These deletions are only safe once context-simple's new semantics are in.

Summary

context-simple's max_tokens used to be a fallback consulted only when a provider published no context window, so in these two e2e pipeline profiles it was inert. It is now a CAP defaulting to None, and leaving the lines in place would clamp both pipelines to 300,000 tokens. This PR deletes session.context.config.max_tokens from profiles/attractor-e2e-pipeline-anthropic.yaml and profiles/attractor-e2e-pipeline-gemini.yaml (2 lines, no other changes).

The Gemini profile is the interesting one. Gemini's provider published no context window at all, so this 300,000 was genuinely live there — the only place in the ecosystem where it was. provider-gemini now publishes its real 1,048,576-token window, so removing the line lets that pipeline use the whole model instead of 29% of it, which is the point.

Verification checklist

  • nlspec evidence: N/A — no spec section governs a bundle profile's max_tokens value. This diff changes no dispatch, admission, or event behavior; it removes two inert config lines from e2e profiles.
  • Unit tests pass — see evidence below.
  • Live pipeline run exercising changed code path — N/A for engine.py/handler code (none touched). The changed path is the profile-to-context-manager seam, exercised instead by the cross-repo DTU run below.
  • AGENTS.md reviewed; repo-specific gates met
  • Backward-compat path unchanged — the Anthropic profile's effective budget is unchanged (its provider already published a window below 300,000); the Gemini profile's budget intentionally rises, which is the stated purpose.
  • Observable contract / specs/EXTENSIONS.md: no entry needed — config-value removal in two e2e profiles, no observable contract change for pipeline authors or downstream consumers.
  • Doc claim about code behavior: none added or changed in this diff — no number, default, vocabulary, or contract claim ships here, so no guard test is needed.
  • Pre-publication leak review: N/A — no new public content class. The diff is two deleted YAML lines; no new directory, artifact type, evidence-bearing doc, or fixture corpus.
  • PR body includes verification evidence, not just "tests pass"
  • CI is green before merge — confirmed at merge time by the merger on head 48e000a5: CI Gate (all checks passed) reports pass, alongside DOT Render Gate, Opinionated Guards, Unit Tests (py3.11 + py3.13) and license/cla — 6/6 green, none bypassed.

Verification evidence

Both files were re-parsed after the edit. The first attempt removed the orchestrator's config: key instead of the context manager's, and the parse check caught it — the error was fixed before commit, and this is why the parse check exists.

Repo suite:

268 passed, 2 skipped

Cross-repo DTU validation, instance context-overflow-fix-20260915, all 8 target behaviors PASS. The headline result is exactly the one this PR's Gemini profile depends on: a Gemini session's effective budget goes from 200,000 (the old fallback) to 1,011,712 (its real published window), a 5.1x increase. Also: max_tokens: 500000 caps to exactly 500,000; a cap above the model window is a no-op; and overrides.context-simple.config in settings.yaml now reaches session.context through the CLI's own resolve_bundle_config.

Notes for reviewers

This is part 5 of a coordinated five-repo change and must merge after context-simple, provider-gemini and app-cli. If merged before context-simple, the deletion restores the old inert-fallback behavior and the Gemini pipeline keeps the 200,000 fallback — not harmful, but the intended gain does not land until the ordering holds. Cross-links to the other four PRs are in a follow-up comment.

Coordinated five-repo change

max_tokens in context-simple was documented as "Maximum context size" but implemented as a fallback consulted only when a provider published no context window. Orchestrators always pass a provider, so the knob was silently dead in production. It is now a real cap (default None = no cap); a new max_tokens_fallback (default 200,000) took over the fallback role.

Merge order matters

Merge FIRST (independent of each other, any order):

  1. amplifier-module-context-simple — feat/max-tokens-cap
  2. amplifier-module-provider-gemini — feat/publish-context-window
  3. amplifier-app-cli — feat/session-module-config-overrides

Merge AFTER those three:
4. amplifier-foundation — chore/drop-dead-max-tokens
5. amplifier-bundle-attractor — chore/drop-dead-max-tokens

Rationale: the two config repos delete max_tokens lines that only become safe-to-delete once context-simple's new semantics are in.

Cross-repo verification

DTU instance context-overflow-fix-20260915, all 8 target behaviors PASS:

  • A Gemini session's effective budget goes from 200,000 (the old fallback) to 1,011,712 (its real published window) — a 5.1x increase.
  • max_tokens: 500000 caps to exactly 500,000.
  • A cap above the model window is a no-op.
  • settings.yaml overrides.context-simple.config now reaches session.context through the CLI's own resolve_bundle_config.
Repo Result
context-simple 166 passed, 1 xfailed
provider-gemini 330 passed; 1 pre-existing failure (test_image_vision_integration_with_real_api, needs a live GOOGLE_API_KEY, fails on main too)
app-cli 2335 passed, 2 skipped, 1 xfailed
amplifier-foundation 1938 passed, 3 skipped; 1 pre-existing failure (test_grpc_adapter_main.py::TestVerifyModuleType::test_non_isinstance_object_with_mount_passes, verified failing on main before the change)
amplifier-bundle-attractor 268 passed, 2 skipped

context-simple's `max_tokens` used to be a fallback consulted only when a
provider published no context window, so in these profiles it was inert.
It is now a CAP defaulting to None, and leaving the lines in place would
clamp both pipelines to 300,000 tokens.

The Gemini profile is the interesting one. Gemini's provider published no
context window at all, so this 300,000 was genuinely live there -- the only
place in the ecosystem where it was. provider-gemini now publishes its real
1,048,576-token window, so removing the line lets that pipeline use the whole
model instead of 29% of it, which is the point.

Both files re-parsed after the edit; the first attempt removed the
orchestrator's `config:` key instead of the context manager's, and the parse
check caught it.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach

Copy link
Copy Markdown
Collaborator Author

The five PRs in this coordinated change

Merge FIRST — independent of each other, any order:

  1. context-simple — feat: max_tokens becomes a real cap; max_tokens_fallback takes over the fallback amplifier-module-context-simple#42
  2. provider-gemini — fix: publish the real context window; refresh rates; purge the stale 8,192 amplifier-module-provider-gemini#48
  3. app-cli — fix: config overrides reach session.context and session.orchestrator amplifier-app-cli#342

Merge AFTER those three:
4. amplifier-foundation — microsoft/amplifier-foundation#388
5. amplifier-bundle-attractor — #359

The two config repos (4, 5) delete max_tokens lines that only become safe-to-delete once context-simple's new cap semantics (1) are in.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants