Skip to content

fix: publish the real context window; refresh rates; purge the stale 8,192 - #48

Merged
Brian Krabach (bkrabach) merged 1 commit into
mainfrom
feat/publish-context-window
Sep 15, 2026
Merged

Brian Krabach (bkrabach) merged 1 commit into
mainfrom
feat/publish-context-window

Conversation

@bkrabach

Copy link
Copy Markdown
Collaborator

What changed

get_info() published no context_window and no max_output_tokens, so context-simple's _calculate_budget fell through to its conservative fallback on every Gemini session: a 200,000-token budget (300,000 where a bundle set one) against models that accept 1,048,576. Roughly 70% of the window went unused, silently — an undersized budget does not error, it just compacts early.

The window was never unknown. list_models() has always read input_token_limit off the live API; the number simply never reached the budget path.

  • New _capabilities.py: per-model context_window / max_output_tokens, with longest-prefix matching so dated aliases resolve to their base model, and a documented default for ids newer than this file. list_models() remains the live source of truth; this answers only the synchronous question get_info() cannot make a network call to answer.
  • get_info() now publishes both keys context-simple reads, for self.default_model — it previously reported a hardcoded "gemini-3.7-flash" no matter what the instance was configured with.
  • The 8,192 output cap is gone from source and README. It predates every model this provider serves; all current models accept 65,536. list_models() also stopped clamping the API's own reported limit to it.

Costs re-verified 2026-09-15 against the pricing page (last updated 2026-09-11)

  • gemini-3.7-flash, the provider's own default model, was missing from the rate table entirely — every call on it reported cost None.
  • gemini-3.8-flash (GA 2026-09-02) and gemini-3.5-flash-lite added.
  • gemini-3.6-flash carried 1.50/7.50 — the post-promo rate — as if current, over-reporting by 2x.
  • 3.6/3.7/3.8 Flash are on promotional pricing through 2026-12-31, after which both rates double. Those entries now carry promo_until plus a post_promo block and compute_cost() selects by date, injectable via today= so both sides of the boundary are testable. A hardcoded promo rate would have silently under-reported by 2x from 2027-01-01 — invisible, because a wrong cost looks exactly like a right one.
  • gemini-3-pro-preview removed: Google shut it down 2026-03-09. It has no published rate, so any number here would be invented; None is the honest answer.

How to verify

New test_published_context_window.py (21 tests) pins the contract at the seam that consumes it. Existing cost/config tests updated to the new rates and limits.

330 passed

Two pre-existing failures/lints are untouched and unrelated, both present on main: test_image_vision_integration_with_real_api needs a live GOOGLE_API_KEY, and ruff F401 in test_streaming.py.

Breaking changes

Cost lookups for gemini-3-pro-preview now return None rather than a rate (the model is shut down). Reported costs for gemini-3.6-flash drop 2x to the actual current promotional rate.

Coordinated five-repo change

max_tokens in context-simple was documented as "Maximum context size" but implemented as a fallback consulted only when a provider published no context window. Orchestrators always pass a provider, so the knob was silently dead in production. It is now a real cap (default None = no cap); a new max_tokens_fallback (default 200,000) took over the fallback role.

Merge order matters

Merge FIRST (independent of each other, any order):

  1. amplifier-module-context-simple — feat/max-tokens-cap
  2. amplifier-module-provider-gemini — feat/publish-context-window
  3. amplifier-app-cli — feat/session-module-config-overrides

Merge AFTER those three:
4. amplifier-foundation — chore/drop-dead-max-tokens
5. amplifier-bundle-attractor — chore/drop-dead-max-tokens

Rationale: the two config repos delete max_tokens lines that only become safe-to-delete once context-simple's new semantics are in.

Cross-repo verification

DTU instance context-overflow-fix-20260915, all 8 target behaviors PASS:

  • A Gemini session's effective budget goes from 200,000 (the old fallback) to 1,011,712 (its real published window) — a 5.1x increase.
  • max_tokens: 500000 caps to exactly 500,000.
  • A cap above the model window is a no-op.
  • settings.yaml overrides.context-simple.config now reaches session.context through the CLI's own resolve_bundle_config.
Repo Result
context-simple 166 passed, 1 xfailed
provider-gemini 330 passed; 1 pre-existing failure (test_image_vision_integration_with_real_api, needs a live GOOGLE_API_KEY, fails on main too)
app-cli 2335 passed, 2 skipped, 1 xfailed
amplifier-foundation 1938 passed, 3 skipped; 1 pre-existing failure (test_grpc_adapter_main.py::TestVerifyModuleType::test_non_isinstance_object_with_mount_passes, verified failing on main before the change)
amplifier-bundle-attractor 268 passed, 2 skipped

…8,192

get_info() published no context_window and no max_output_tokens, so
context-simple's _calculate_budget fell through to its conservative fallback
on EVERY Gemini session: a 200,000-token budget (300,000 where a bundle set
one) against models that accept 1,048,576. Roughly 70% of the window went
unused, silently -- an undersized budget does not error, it just compacts
early.

The window was never unknown. list_models() has always read
input_token_limit off the live API; the number simply never reached the
budget path.

- New _capabilities.py: per-model context_window / max_output_tokens, with
  longest-prefix matching so dated aliases resolve to their base model, and a
  documented default for ids newer than this file. list_models() remains the
  live source of truth; this answers only the synchronous question get_info()
  cannot make a network call to answer.
- get_info() now publishes both keys context-simple reads, for
  self.default_model -- it previously reported a hardcoded "gemini-3.7-flash"
  no matter what the instance was configured with.
- The 8,192 output cap is gone from source and README. It predates every model
  this provider serves; all current models accept 65,536. list_models() also
  stopped clamping the API's own reported limit to it.

Costs re-verified 2026-09-15 against the pricing page (last updated
2026-09-11):

- gemini-3.7-flash, the provider's OWN DEFAULT MODEL, was missing from the
  rate table entirely -- every call on it reported cost None.
- gemini-3.8-flash (GA 2026-09-02) and gemini-3.5-flash-lite added.
- gemini-3.6-flash carried 1.50/7.50 -- the POST-promo rate -- as if current,
  over-reporting by 2x.
- 3.6/3.7/3.8 Flash are on promotional pricing through 2026-12-31, after which
  both rates double. Those entries now carry promo_until plus a post_promo
  block and compute_cost() selects by date, injectable via today= so both
  sides of the boundary are testable. A hardcoded promo rate would have
  silently under-reported by 2x from 2027-01-01 -- invisible, because a wrong
  cost looks exactly like a right one.
- gemini-3-pro-preview removed: Google shut it down 2026-03-09. It has no
  published rate, so any number here would be invented; None is the honest
  answer.

Tests: new test_published_context_window.py (21) pins the contract at the seam
that consumes it. Existing cost/config tests updated to the new rates and
limits. 330 passed. Two pre-existing failures/lints are untouched and
unrelated: test_image_vision_integration_with_real_api needs a live
GOOGLE_API_KEY, and ruff F401 in test_streaming.py -- both present on main.

Generated with Amplifier

Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
@bkrabach

Copy link
Copy Markdown
Collaborator Author

The five PRs in this coordinated change

Merge FIRST — independent of each other, any order:

  1. context-simple — feat: max_tokens becomes a real cap; max_tokens_fallback takes over the fallback amplifier-module-context-simple#42
  2. provider-gemini — fix: publish the real context window; refresh rates; purge the stale 8,192 #48
  3. app-cli — fix: config overrides reach session.context and session.orchestrator amplifier-app-cli#342

Merge AFTER those three:
4. amplifier-foundation — microsoft/amplifier-foundation#388
5. amplifier-bundle-attractor — microsoft/amplifier-bundle-attractor#359

The two config repos (4, 5) delete max_tokens lines that only become safe-to-delete once context-simple's new cap semantics (1) are in.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants