fix: publish the real context window; refresh rates; purge the stale 8,192 - #48
Merged
Merged
Conversation
…8,192 get_info() published no context_window and no max_output_tokens, so context-simple's _calculate_budget fell through to its conservative fallback on EVERY Gemini session: a 200,000-token budget (300,000 where a bundle set one) against models that accept 1,048,576. Roughly 70% of the window went unused, silently -- an undersized budget does not error, it just compacts early. The window was never unknown. list_models() has always read input_token_limit off the live API; the number simply never reached the budget path. - New _capabilities.py: per-model context_window / max_output_tokens, with longest-prefix matching so dated aliases resolve to their base model, and a documented default for ids newer than this file. list_models() remains the live source of truth; this answers only the synchronous question get_info() cannot make a network call to answer. - get_info() now publishes both keys context-simple reads, for self.default_model -- it previously reported a hardcoded "gemini-3.7-flash" no matter what the instance was configured with. - The 8,192 output cap is gone from source and README. It predates every model this provider serves; all current models accept 65,536. list_models() also stopped clamping the API's own reported limit to it. Costs re-verified 2026-09-15 against the pricing page (last updated 2026-09-11): - gemini-3.7-flash, the provider's OWN DEFAULT MODEL, was missing from the rate table entirely -- every call on it reported cost None. - gemini-3.8-flash (GA 2026-09-02) and gemini-3.5-flash-lite added. - gemini-3.6-flash carried 1.50/7.50 -- the POST-promo rate -- as if current, over-reporting by 2x. - 3.6/3.7/3.8 Flash are on promotional pricing through 2026-12-31, after which both rates double. Those entries now carry promo_until plus a post_promo block and compute_cost() selects by date, injectable via today= so both sides of the boundary are testable. A hardcoded promo rate would have silently under-reported by 2x from 2027-01-01 -- invisible, because a wrong cost looks exactly like a right one. - gemini-3-pro-preview removed: Google shut it down 2026-03-09. It has no published rate, so any number here would be invented; None is the honest answer. Tests: new test_published_context_window.py (21) pins the contract at the seam that consumes it. Existing cost/config tests updated to the new rates and limits. 330 passed. Two pre-existing failures/lints are untouched and unrelated: test_image_vision_integration_with_real_api needs a live GOOGLE_API_KEY, and ruff F401 in test_streaming.py -- both present on main. Generated with Amplifier Co-Authored-By: Amplifier <240397093+microsoft-amplifier@users.noreply.github.com>
Collaborator
Author
The five PRs in this coordinated changeMerge FIRST — independent of each other, any order:
Merge AFTER those three: The two config repos (4, 5) delete |
This was referenced Sep 15, 2026
fix: config overrides reach session.context and session.orchestrator
microsoft/amplifier-app-cli#342
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
get_info()published nocontext_windowand nomax_output_tokens, so context-simple's_calculate_budgetfell through to its conservative fallback on every Gemini session: a 200,000-token budget (300,000 where a bundle set one) against models that accept 1,048,576. Roughly 70% of the window went unused, silently — an undersized budget does not error, it just compacts early.The window was never unknown.
list_models()has always readinput_token_limitoff the live API; the number simply never reached the budget path._capabilities.py: per-modelcontext_window/max_output_tokens, with longest-prefix matching so dated aliases resolve to their base model, and a documented default for ids newer than this file.list_models()remains the live source of truth; this answers only the synchronous questionget_info()cannot make a network call to answer.get_info()now publishes both keys context-simple reads, forself.default_model— it previously reported a hardcoded"gemini-3.7-flash"no matter what the instance was configured with.list_models()also stopped clamping the API's own reported limit to it.Costs re-verified 2026-09-15 against the pricing page (last updated 2026-09-11)
gemini-3.7-flash, the provider's own default model, was missing from the rate table entirely — every call on it reported costNone.gemini-3.8-flash(GA 2026-09-02) andgemini-3.5-flash-liteadded.gemini-3.6-flashcarried1.50/7.50— the post-promo rate — as if current, over-reporting by 2x.promo_untilplus apost_promoblock andcompute_cost()selects by date, injectable viatoday=so both sides of the boundary are testable. A hardcoded promo rate would have silently under-reported by 2x from 2027-01-01 — invisible, because a wrong cost looks exactly like a right one.gemini-3-pro-previewremoved: Google shut it down 2026-03-09. It has no published rate, so any number here would be invented;Noneis the honest answer.How to verify
New
test_published_context_window.py(21 tests) pins the contract at the seam that consumes it. Existing cost/config tests updated to the new rates and limits.Two pre-existing failures/lints are untouched and unrelated, both present on
main:test_image_vision_integration_with_real_apineeds a liveGOOGLE_API_KEY, and ruff F401 intest_streaming.py.Breaking changes
Cost lookups for
gemini-3-pro-previewnow returnNonerather than a rate (the model is shut down). Reported costs forgemini-3.6-flashdrop 2x to the actual current promotional rate.Coordinated five-repo change
max_tokensincontext-simplewas documented as "Maximum context size" but implemented as a fallback consulted only when a provider published no context window. Orchestrators always pass a provider, so the knob was silently dead in production. It is now a real cap (defaultNone= no cap); a newmax_tokens_fallback(default 200,000) took over the fallback role.Merge order matters
Merge FIRST (independent of each other, any order):
amplifier-module-context-simple—feat/max-tokens-capamplifier-module-provider-gemini—feat/publish-context-windowamplifier-app-cli—feat/session-module-config-overridesMerge AFTER those three:
4.
amplifier-foundation—chore/drop-dead-max-tokens5.
amplifier-bundle-attractor—chore/drop-dead-max-tokensRationale: the two config repos delete
max_tokenslines that only become safe-to-delete once context-simple's new semantics are in.Cross-repo verification
DTU instance
context-overflow-fix-20260915, all 8 target behaviors PASS:max_tokens: 500000caps to exactly 500,000.settings.yamloverrides.context-simple.confignow reachessession.contextthrough the CLI's ownresolve_bundle_config.test_image_vision_integration_with_real_api, needs a liveGOOGLE_API_KEY, fails onmaintoo)test_grpc_adapter_main.py::TestVerifyModuleType::test_non_isinstance_object_with_mount_passes, verified failing onmainbefore the change)