Google Gemini model integration for Amplifier via Google AI API.
The simplest way to use Gemini with Amplifier. Once installed, Amplifier will automatically discover it.
-
Set your API key:
export GOOGLE_API_KEY="your-api-key-here"
Get your API key from Google AI Studio.
-
Add the module:
amplifier module add provider-gemini --source git+https://github.com/microsoft/amplifier-module-provider-gemini@main --global
-
Use Gemini as your provider:
amplifier provider use gemini --global
-
Start using it:
amplifier run "Hello from Gemini!"
That's it! The module is now available for all your projects. Use --project instead of --global to install for just the current project.
For more control over configuration or to compose with other capabilities, use a bundle:
-
Set your API key:
export GOOGLE_API_KEY="your-api-key-here"
Get your API key from Google AI Studio.
-
Create a bundle in your project or home directory (e.g.,
gemini-bundle/bundle.md):--- bundle: name: gemini-dev version: 1.0.0 description: Gemini provider with full 1M context includes: - bundle: foundation session: context: config: max_tokens: 1048576 # Full 1M input context providers: - module: provider-gemini source: git+https://github.com/microsoft/amplifier-module-provider-gemini@main config: default_model: gemini-3.8-flash max_output_tokens: 65536 # Full 65K output capacity temperature: 0.7 priority: 50 # Lower number = higher priority (beats default 100) --- # Gemini Development Bundle This bundle configures Gemini with full context windows and includes foundation capabilities. ## Available Models - **Gemini 3.8 Flash** - `gemini-3.8-flash` - Default Flash model; explicit pins remain unchanged - **Gemini 3.5 Flash / Flash-Lite** - `gemini-3.5-flash` / `gemini-3.5-flash-lite` - Legacy Flash generation - **Gemini 2.5 Flash / Flash-Lite / Pro** - `gemini-2.5-flash` / `gemini-2.5-flash-lite` / `gemini-2.5-pro` - Two generations back; still served
-
Use it:
amplifier run --bundle ./gemini-bundle "Hello from Gemini!"
Offline provider contracts run without real credentials: scoped fixtures mount
with a nonfunctional key and mock only the SDK async model pager. Run
uv run pytest -q -m "not live" for all offline checks.
- Python 3.11+
- UV - Fast Python package manager
# macOS/Linux/WSL
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"Provides access to Google's Gemini models as an LLM provider for Amplifier with 1M token context windows and extended thinking capabilities.
Module Type: Provider
Mount Point: providers
Entry Point: amplifier_module_provider_gemini:mount
Current support: Text generation, tool calling, and thinking. Multimodal capabilities (images, video, audio) are not yet implemented.
Model availability and naming change frequently -- this list reflects what was verified live against list_models() on 2026-08-29. Always prefer amplifier provider models gemini (or your account's actual list_models() result) over this table for what your key can currently use.
gemini-3.8-flash- Default model for this provider. Selected from the October 2026 catalog refresh; explicit pins remain unchanged. Supported thinking levels follow the provider's model-specific capability rules.gemini-3.7-flash- Previous default. Usesthinking_level(low/medium/high -- nominimal, verified in August).gemini-3.5-flash/gemini-3.5-flash-lite- Documented by Google as legacy relative to 3.7. Usesthinking_level(full minimal/low/medium/high range).- Other
gemini-3.*preview/dated ids (e.g.gemini-3.1-flash-lite-preview,gemini-3-pro-image-preview) come and go -- this provider assumes anygemini-3.*id supportsthinking_levelwith the full range unless proven otherwise by a live 400.
gemini-2.5-flash- Best price-performance for large-scale processing (1M context, 65K max output)gemini-2.5-pro- State-of-the-art thinking model for complex reasoning (1M context, 65K max output)gemini-2.5-flash-lite- Fastest model optimized for cost-efficiency (1M context, 65K max output)
Verified live (2026-08-29): these three REJECT thinking_level outright ("Thinking level is not supported for this model") -- this provider automatically falls back to the legacy thinking_budget control for them. See Thinking/Reasoning below.
Shut down by Google -- no longer served. Do not configure gemini-2.0-flash / gemini-2.0-flash-lite as your model.
Note: Image/video/audio models not listed as the provider doesn't support multimodal capabilities yet.
[[providers]]
module = "provider-gemini"
name = "gemini"
config = {
default_model = "gemini-3.8-flash",
max_output_tokens = 65536,
temperature = 0.7,
}Every key below corresponds either to a real parameter in Google's google.genai.types.GenerateContentConfig (linked as "API: field_name") or is Amplifier-only glue with no Google equivalent (marked "Amplifier-only").
| Parameter | Type | Default | API param / origin | Description |
|---|---|---|---|---|
api_key |
string | env: GOOGLE_API_KEY or GEMINI_API_KEY |
Amplifier-only | Google AI API key. Env vars match the official SDK's own resolution (GOOGLE_API_KEY wins if both are set). |
default_model |
string | gemini-3.8-flash |
Amplifier-only | Default model to use when a request doesn't override it. |
max_output_tokens |
int | model default (65,536 for every current model) | API: max_output_tokens |
Maximum output tokens. Renamed from max_tokens to match Google's own parameter name -- max_tokens still works as a deprecated alias (one-shot warning; max_output_tokens always wins if both are set). |
max_tokens |
int | -- | (deprecated alias) | Old name for max_output_tokens. Prefer the new name in new configs. |
temperature |
float | 0.7 | API: temperature |
Sampling temperature (0.0-2.0 per Google's docs; this provider does not clamp the range itself). |
timeout |
float or null | null | Provider and SDK | Optional API call / stream idle timeout in seconds; default waits for completion. |
close_timeout |
float | 5.0 | Amplifier-only | Hard bound on closing the genai client at session teardown, seconds. Both of the SDK client's httpx transports are closed (client.aio.aclose() for the async one this provider actually uses, and the synchronous client.close(), offloaded to a worker thread so it never blocks the event loop); neither has a deadline of its own and either can block forever on a half-closed (CLOSE-WAIT) connection. On timeout the client is abandoned with a WARNING and the sockets are reclaimed at process exit. |
priority |
int | 100 | Amplifier-only | Provider selection priority (lower = preferred). Read by the orchestrator's provider-selection logic, not by this module's own request-building code. |
raw |
bool | false | Amplifier-only | Enable raw API request/response capture on the llm:request / llm:response events (non-streaming path only). |
use_streaming |
bool | true | Amplifier-only | Use generate_content_stream instead of a single blocking generate_content call. Per-request override: request.metadata["stream"] = False. |
max_retries |
int | 5 | Amplifier-only | Max retry attempts on transient failures (5xx, timeouts, rate limits, Cloudflare/CDN challenges). |
min_retry_delay |
float | 1.0 | Amplifier-only | Initial retry backoff delay, in seconds. |
max_retry_delay |
float | 60.0 | Amplifier-only | Maximum retry backoff delay, in seconds. |
retry_jitter |
bool | true | Amplifier-only | Add random jitter to retry backoff delays. |
max_concurrent_requests |
int | 5 | Amplifier-only | Process-wide concurrency limit shared across all provider instances (parent + delegated sessions). 0 disables the limit. |
extra_request_params |
dict | {} |
(passthrough) | Settings-only escape hatch -- see extra_request_params below. Never an interactive config field; owner-beware. |
Boolean and numeric values above tolerate string input ("true"/"false", "600", etc. -- as written by the app-cli wizard or hand-edited YAML) and warn-and-default rather than crash on anything unparseable. Unrecognized config keys log a warning at mount time (with a "did you mean" suggestion for likely typos); they never silently do nothing without a signal.
Removed in this revision:
debug,raw_debug, anddebug_truncate_lengthwere documented here in earlier README revisions but were never actually implemented by this provider (no code path reads them). If your config still sets them, they are now flagged with a specific "this key is inert" warning at mount time instead of being silently ignored. Useraw: truefor this provider's actual raw-I/O capture onllm:request/llm:responseevents.
extra_request_params merges arbitrary fields directly into the GenerateContentConfig this provider builds for every request -- reaching Google API parameters this provider doesn't otherwise expose as a first-class config key: safety_settings, top_p, top_k, seed, stop_sequences, presence_penalty, frequency_penalty, response_mime_type, labels, and anything else google.genai.types.GenerateContentConfig defines.
providers:
- module: provider-gemini
config:
default_model: gemini-3.8-flash
extra_request_params:
top_p: 0.95
safety_settings:
- category: HARM_CATEGORY_DANGEROUS_CONTENT
threshold: BLOCK_ONLY_HIGHContract:
- Merged LAST, after every value this provider computes itself (temperature,
max_output_tokens,thinking_config, tools). Yourextra_request_paramsvalue always wins. - Overriding a value this provider had already set logs a warning naming the field, this provider's computed value, and your override -- a silent production override never happens.
thinking_configis the nested exception: a mapping or SDKThinkingConfigmerges only the fields explicitly supplied into the provider's computed typed config. An empty mapping orThinkingConfig()is a no-op; explicitfalseis preserved, while explicitnullclears that field. Both SDK spellings work (thinking_level/thinkingLevel,thinking_budget/thinkingBudget, andinclude_thoughts/includeThoughts). Additive and equal-value updates are quiet; conflicting replacements produce one warning listing the changed field names, without dumping their values.- Inside that merge,
nullexplicitly clears its field. A non-nullthinking_levelclears an inheritedthinking_budget, and vice versa, because Gemini rejects requests containing both. Supplying both non-null fields together is invalid and raisesValueErrorbefore the request is sent. Top-levelthinking_config: nullremains a deliberate whole-field override. - An unrecognized field name (not a real
GenerateContentConfigfield) logs a warning and is skipped -- it never crashes the provider mount. - Settings-only, deliberately not a
ConfigField: it will never appear in the interactiveamplifier init/amplifier provider usewizard. Set it directly insettings.yamlor a bundle's config block. If your tooling round-trips provider config (e.g. re-serializing settings), this key passes through unchanged like any other dict value -- there is no special handling on this provider's side beyond the merge described above. - Safety filters default OFF on Gemini 2.5/3.x models (verified against ai.google.dev) -- this provider never injects a
safety_settingsdefault of its own. If you want filtering, set it explicitly viaextra_request_params.safety_settings.
The provider supports both environment variables (matching the official Google GenAI SDK):
# Either of these works (GOOGLE_API_KEY takes precedence if both are set)
export GEMINI_API_KEY="your-api-key-here"
# or
export GOOGLE_API_KEY="your-api-key-here"Get your API key from Google AI Studio.
# In amplifier configuration
[provider]
name = "gemini"
default_model = "gemini-3.8-flash"For advanced configuration, add Gemini to any bundle. These examples show the YAML configuration section (the frontmatter between --- markers in your bundle.md file):
Basic Configuration:
providers:
- module: provider-gemini
source: git+https://github.com/microsoft/amplifier-module-provider-gemini@main
config:
default_model: gemini-3.8-flash
max_output_tokens: 65536 # Use full 65K output capacity
temperature: 0.7
priority: 50 # IMPORTANT: Lower number = higher priority (beats default 100)Balanced (1M context, cost-effective):
providers:
- module: provider-gemini
source: git+https://github.com/microsoft/amplifier-module-provider-gemini@main
config:
default_model: gemini-3.8-flash
max_output_tokens: 65536 # Full 65K output capacity
priority: 50 # Lower number = higher priorityThinking (complex reasoning with full 1M context):
session:
context:
config:
max_tokens: 1048576 # Full 1M input context
orchestrator:
module: loop-streaming
source: git+https://github.com/microsoft/amplifier-module-loop-streaming@main
config:
extended_thinking: true # Show thinking content
providers:
- module: provider-gemini
source: git+https://github.com/microsoft/amplifier-module-provider-gemini@main
config:
default_model: gemini-3.8-flash
max_output_tokens: 65536 # Full 65K output capacity
temperature: 1.0
priority: 50 # Lower number = higher priorityFast (simple queries, low cost):
providers:
- module: provider-gemini
source: git+https://github.com/microsoft/amplifier-module-provider-gemini@main
config:
default_model: gemini-3.5-flash-lite
max_output_tokens: 65536 # Full 65K output capacity
temperature: 0.5
priority: 50 # Lower number = higher priority- Text Generation - Single and multi-turn conversations
- Tool/Function Calling - OpenAPI schema format
- Extended Thinking - Reasoning via
thinking_level(current models) orthinking_budget(legacy models) - Streaming Support - Incremental response generation
- 1M Token Context - Process extremely large inputs (Flash models)
- Message Validation - Defense-in-depth error checking
Google's thinking control surface changed generations: thinking_budget (an approximate output-token budget) is the legacy control; thinking_level (an enum: minimal/low/medium/high) is the current control -- and, as of this revision, the only control some models accept at all. Sending both on one request is rejected by the API with a 400.
This provider maps Amplifier's portable reasoning_effort request field to thinking_level automatically, per-model, clamped against a small maintained support table:
| Model | Supported thinking_level values |
Notes |
|---|---|---|
gemini-3.7-flash |
low, medium, high |
Rejects minimal (verified live: "Thinking level MINIMAL is not supported for this model"). Google's own default (when omitted) is medium. |
gemini-3.5-flash, gemini-3.5-flash-lite |
minimal, low, medium, high |
Full range. |
Other gemini-3.* ids |
minimal, low, medium, high (assumed) |
Not individually verified -- assumed full range until a live 400 proves narrower. |
gemini-2.5-flash, gemini-2.5-pro, gemini-2.5-flash-lite |
none | Rejects thinking_level entirely (verified live: "Thinking level is not supported for this model"). Falls back to the legacy thinking_budget mapping below. |
gemini-2.0-* |
none (shut down) | Not servable at all. |
reasoning_effort -> thinking_level mapping (for models that support it):
reasoning_effort |
Target level | If not supported by the model |
|---|---|---|
none |
(no level sent) | Uses minimal if the model supports it, otherwise the model's own default (Gemini 3.x cannot disable thinking at all -- verified live) |
minimal |
minimal |
Clamped up to the nearest supported level (e.g. low on gemini-3.7-flash), logged at INFO |
low |
low |
Clamped as above |
medium |
medium |
Clamped as above |
high, xhigh, max |
high |
Gemini has no level above high |
For models with no thinking_level support (the gemini-2.5-* family), reasoning_effort instead maps to the legacy numeric thinking_budget: none -> 0 (disabled), minimal/low -> 4096, medium/high/xhigh/max -> -1 (dynamic, model decides).
An explicit thinking_budget passed via request.metadata["thinking_budget"] or a provider **kwargs override always wins outright and is sent alone -- never combined with thinking_level on the same request.
To display thinking output, configure your orchestrator (not the provider):
session:
orchestrator:
module: loop-streaming # Required for thinking display
source: git+https://github.com/microsoft/amplifier-module-loop-streaming@main
config:
extended_thinking: true # Show thinking content to userNote: The provider captures thinking from the API automatically (include_thoughts defaults to true). The orchestrator's extended_thinking: true config controls whether it's displayed. Without this config, thinking still happens but isn't shown to the user.
Thought signatures: Gemini 2.5+ models attach an opaque thought_signature to parts that follow a thinking burst. Because this provider is stateless full-resend (the entire conversation is rebuilt from stored history and resent on every turn, with no server-side session), these signatures must be captured and replayed unmodified or the API returns FinishReason MISSING_THOUGHT_SIGNATURE. This provider captures and echoes signatures on text, thinking, and tool-call parts alike, encoded as base64 so the value survives any JSON serialization the conversation history passes through (e.g. session persistence, event logging).
Functions are declared using OpenAPI schema format:
tools = [{
"name": "get_weather",
"description": "Get weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City name"}
},
"required": ["location"]
}
}]The provider handles tool call marshaling and response integration automatically.
The provider implements automatic repair for incomplete tool call sequences:
The Problem: If tool results are missing from conversation history (due to context compaction bugs, parsing errors, or state corruption), the Gemini API rejects the entire request, breaking the user's session.
The Solution: The provider automatically detects and repairs missing tool_results by injecting synthetic results:
- Repair before API call - Detects missing tool_results and injects synthetic ones
- Make failures visible - Synthetic results contain
[SYSTEM ERROR: Tool result missing]messages - Maintain conversation validity - API accepts repaired messages, session continues
- Enable recovery - LLM acknowledges error and can ask user to retry
- Provide observability - Emits
provider:tool_sequence_repairedevent with repair details
Example:
# Conversation with missing tool result
messages = [
{
"role": "assistant",
"content": [
{"type": "tool_call", "id": "gemini_call_abc123", "name": "get_weather", "input": {...}}
]
},
# MISSING: {"role": "tool", "tool_call_id": "gemini_call_abc123", "content": "..."}
{"role": "user", "content": "Thanks"}
]
# Provider repairs by injecting synthetic result:
{
"role": "tool",
"tool_call_id": "gemini_call_abc123",
"name": "get_weather",
"content": "[SYSTEM ERROR: Tool result missing from conversation history]\n\nTool: get_weather\nCall ID: gemini_call_abc123\n\nThis indicates the tool result was lost after execution.\nLikely causes: context compaction bug, message parsing error, or state corruption.\n\nThe tool may have executed successfully, but the result was lost.\nPlease acknowledge this error and offer to retry the operation."
}This is a defense-in-depth safety net. The orchestrator should handle tool execution errors at runtime, so this repair only triggers when results go missing due to bugs in context management.
The Gemini API does not provide tool call IDs (unlike Anthropic and OpenAI). The provider generates synthetic IDs using the format gemini_call_{uuid} to maintain compatibility with Amplifier's tool protocol.
Impact: Tool call IDs are unique and functional but not provided by the API itself. This is transparent to users but documented for debugging purposes.
The provider implements text generation, tool calling, and thinking support. Multimodal capabilities (images, video, audio) are not yet supported.
Verified live: thinking_budget=0 on gemini-3.7-flash still produced thinking tokens, and there is no "off" thinking_level. Thinking is mandatory for Gemini 3.x models regardless of what this provider sends. reasoning_effort="none" on these models falls back to the model's own default thinking amount rather than actually disabling it -- this is a vendor limitation, not something this provider can work around.
For a configured, limit-tabled gemini-* Developer API model, request_budget() can call
POST /v1beta/models/{model}:countTokens with the documented nested
generateContentRequest body. It uses the same provider-owned request plan as
generation and includes converted system/developer messages, inline images,
thought signatures, function declarations/results, toolConfig,
safetySettings, generationConfig, and cachedContent.
The nested generateContentRequest carries its own REQUIRED model field
(models/<selected model>), per the REST CountTokensRequest contract. It is
always built from the same selected model as the :countTokens endpoint path,
so the URL and the counted body can never name different models.
When request_budget() returns no decision, it also makes a best-effort
provider:request_budget_unavailable observability event carrying
provider, method, an optional selected model, a fixed reason code, and an
integer http_status when one exists. The reason codes are
unsupported_model, unsupported_route, request_plan_unavailable,
request_projection_unavailable, invalid_output_limit, http_error, and
invalid_response. Exactly one event is emitted per failed budget call, and a
successful count emits none. Nothing else travels with it: no exception text,
response body, request body, URL, credential, or environment value, and a
model ID is included only when it exactly matches a documented ID in the
provider's limit table. Unknown strings, URLs, suffixed aliases and non-string
values are omitted rather than coerced; limit lookup behavior is unchanged.
A failing subscriber cannot turn a count into a product failure; with no hooks
channel at all the same reason is logged as a warning instead.
The counter uses the response's totalTokens as the input measurement. It does
not add cachedContentTokenCount a second time, does not estimate output, and
does not increase the documented raw input limit when a caller lowers
max_output_tokens. A missing/malformed response, HTTP failure, or request
that cannot be publicly serialized returns no budget decision rather than a
partial native count. This includes the Developer-unsupported or
transformer-specific labels, routing/model-selection, response-schema,
speech/image/audio, and Model Armor extras. Cancellation propagates without
starting generation. An unknown model ID is also unavailable: an exact count
cannot justify an inferred input-window admission limit. An extra_request_params
override for tools or system_instruction, and a noncanonical
cached_content value, are likewise unavailable rather than reimplementing
the SDK's private normalization.
This adapter is not Vertex AI support and does not use
client.models.count_tokens: that Python helper cannot send a full Gemini
Developer generateContentRequest. The request-projection wire regression is
written against google-genai 2.23.0, the retained DTU version; no other SDK
version is covered by this exact generation/count wire check. The checked-in
uv.lock still pins 1.46.0 despite this package's pre-existing >=1.56.0
floor; this baseline lock discrepancy is intentionally not upgraded here.
Counting is unavailable when the SDK would route generation away from the
canonical Gemini Developer endpoint: a nonempty GOOGLE_GEMINI_BASE_URL, or
effective Vertex/Enterprise selection through GOOGLE_GENAI_USE_VERTEXAI /
GOOGLE_GENAI_USE_ENTERPRISE, disables both count advertisement and dispatch.
The SDK's 2.23 parsing is followed exactly: true/1 are case-insensitive,
and a present Enterprise variable wins over Vertex even when it is explicitly
false. A client constructed by this provider retains its route decision; an
injected client has no public base-URL guarantee and therefore fails closed.
get_info() reads only these route-selection variables and never constructs
an SDK client, authenticates, or performs I/O.
google-genai>=1.56.0- Official Google AI Python SDK. 1.56.0 is the floor because it's the first release whoseThinkingConfigexposes the fullthinking_levelenum (minimal/low/medium/high) this provider needs -- verified by probing the SDK's own installed types directly: 1.46.0 has nothinking_levelfield at all; 1.51.0 adds it with onlyLOW/HIGH; 1.56.0 completes the four-level enum.httpx>=0.28.1- Async transport for the documented DevelopercountTokensREST endpoint.
Test your local provider changes with the installed amplifier CLI:
# From the provider repository root
cd amplifier-module-provider-gemini
# Add your local provider (use --local for development, not --global)
amplifier module add provider-gemini --source file://. --local
# Set your API key
export GOOGLE_API_KEY="your-api-key-here"
# Now you can use amplifier init and see your local provider in the menu
amplifier init
# Or configure it directly
amplifier provider use gemini --local
# Test your local changes
amplifier run "Hello, testing local Gemini provider!"
# List modules to verify your local provider is registered
amplifier module list -t providerWhen you're done testing:
# Remove the local provider registration
amplifier module remove provider-gemini --localWhy --local instead of --global?
--localregisters the provider only for the current working directory--globalwould affect all your projects (not ideal during development)--projectworks if you want to share with your team
For rapid iteration without the CLI:
cd amplifier-module-provider-gemini
# Install dependencies
uv sync --dev
# Run validation tests (protocol compliance)
uv run pytest
# Test with coverage
uv run pytest --covThese tests validate the provider implements the required protocol without needing the full CLI.
Note
This project is not currently accepting external contributions, but we're actively working toward opening this up. We value community input and look forward to collaborating in the future. For now, feel free to fork and experiment!
Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit Contributor License Agreements.
When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.
This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.
This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.
The default timeout is null: model work waits for completion, explicit cancellation, or a provider/transport error. Set a numeric timeout in seconds to opt into a deadline. Existing cleanup limits are unchanged.