Skip to content

About

Reference implementation of a vLLM provider module for the Amplifier project

Resources

Code of conduct

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Repository files navigation

amplifier-module-provider-vllm

vLLM provider module for Amplifier - Responses API integration for local/self-hosted LLMs.

Overview

This provider module integrates vLLM's OpenAI-compatible Responses API with Amplifier, enabling the use of open-weight models like gpt-oss-20b with full reasoning and tool calling support.

Key Features:

  • Responses API only - Optimized for reasoning models (gpt-oss, etc.)
  • Full reasoning support - Automatic reasoning block separation
  • Tool calling - Complete tool integration via Responses API
  • Local OR remote - Works against a local vLLM with no auth, or any remote / hosted endpoint with Bearer auth
  • OpenAI-compatible - Uses OpenAI SDK under the hood

Installation

# Via uv (recommended)
uv pip install git+https://github.com/microsoft/amplifier-module-provider-vllm@main

# For development
git clone https://github.com/microsoft/amplifier-module-provider-vllm
cd amplifier-module-provider-vllm
uv pip install -e .

Note for GPT-OSS models: Token accounting requires vocab files that are automatically downloaded to ~/.amplifier/cache/vocab/ on first use (requires internet access). If working offline, see troubleshooting section for manual setup.

vLLM Server Setup

This provider requires a running vLLM server. Example setup:

# Start vLLM server (basic)
vllm serve openai/gpt-oss-20b \
  --host 0.0.0.0 \
  --port 8000 \
  --tensor-parallel-size 2

# For production (recommended - full config in /etc/vllm/model.env)
sudo systemctl start vllm

Server requirements:

  • vLLM version: ≥0.10.1 (tested with 0.10.1.1)
  • Responses API: Automatically available (no special flags needed)
  • Model: Any model compatible with vLLM (gpt-oss, Llama, Qwen, etc.)

Configuration

Minimal Configuration

providers:
  - module: provider-vllm
    source: git+https://github.com/microsoft/amplifier-module-provider-vllm@main
    config:
      base_url: "http://192.168.128.5:8000/v1"  # Your vLLM server

Full Configuration

providers:
  - module: provider-vllm
    source: git+https://github.com/microsoft/amplifier-module-provider-vllm@main
    config:
      # Connection
      base_url: "http://192.168.128.5:8000/v1"  # Required: vLLM server URL

      # Model settings
      default_model: "openai/gpt-oss-20b"  # Model name from vLLM
      max_tokens: 4096                      # Max output tokens
      temperature: 0.7                      # Sampling temperature

      # Reasoning
      reasoning: "high"                     # Reasoning effort: minimal|low|medium|high
      reasoning_summary: "detailed"         # Summary verbosity: auto|concise|detailed

      # Context limits (advertised to context managers)
      context_window: 128000                # Model context window in tokens
      max_output_tokens: 32768              # Model max output tokens

      # Advanced
      enable_state: false                   # Server-side state (requires vLLM config)
      truncation: "disabled"                # Fail loud (HTTP 400) instead of silently dropping input (see below)
      timeout: 300.0                        # API timeout (seconds)
      stream_idle_timeout: 300.0            # Max seconds between streamed chunks (see below)
      priority: 100                         # Provider selection priority

      # Debug
      raw: false                            # Attach exact request params to llm:request

debug, raw_debug, and debug_truncate_length are ghost keys from an older version of this README -- they were never read by this provider. Use raw: true to attach the exact request params sent to the server on the llm:request event.

Stream idle timeout

Model requests and streaming reads have no elapsed-time deadline by default. They wait for completion, cancellation, or a real provider/transport failure. Quiet prefill or thinking alone does not establish a failed connection.

Set timeout to opt into a request deadline. Set stream_idle_timeout or VLLM_STREAM_IDLE_TIMEOUT to opt into a limit between streamed chunks, including the first chunk. Both defaults are null; an explicit null config value also turns off an environment-provided idle limit. Explicit numeric limits retain retryable timeout errors. Connection and pool acquisition stay bounded to 5 seconds, and close_timeout continues to bound cleanup.

Context limits

vLLM's /v1/models model cards expose the loaded model's real context length (max_model_len), so context_window is auto-discovered per model the first time list_models() runs — an endpoint serving several models with different limits reports each one accurately instead of a single flat number. max_output_tokens is never discovered this way (vLLM's model cards don't carry it), so it always comes from config/defaults.

Leave context_window unset to auto-detect from the model, the same way num_ctx: 0 behaves in the Ollama provider. Setting it explicitly (config key, or the VLLM_CONTEXT_WINDOW environment variable) overrides discovery, clamped to the server's reported limit so a stale config value can never guarantee a 400. Servers that don't report max_model_len fall back to the configured value or the 128000 default, exactly as before.

This matters for long-context deployments: with the previous hardcoded values (max_output_tokens: 128000), the effective input budget was capped at ~59k tokens even on endpoints that comfortably handle 120k+ token prompts — causing premature compaction thrashing. max_output_tokens here is the advertised model maximum used for budgeting, distinct from max_tokens (the per-request completion cap).

Truncation

BEHAVIOR CHANGE: truncation now defaults to "disabled" (previously "auto"). With truncation: "auto", vLLM's Responses API silently drops input content that overflows the context window: HTTP 200, no error, no warning — the model simply answers from whatever survived. Verified live against a direct vLLM endpoint (glm-5.2, real max_model_len=131072): a ~150,000-token prompt with truncation: "auto" returned HTTP 200 with usage.input_tokens=131056 (~19,000 tokens silently discarded), while the identical prompt with truncation: "disabled" returned a clear HTTP 400 naming the exact limit. OpenAI's own Responses API defaults truncation to "disabled" for the same reason: a caller that cannot detect data loss cannot recover from it.

If your deployment relied on the old auto-truncating behavior (e.g. because an upstream context manager doesn't yet cap prompt size to the model's context window), opt back in explicitly:

config:
  truncation: "auto"

Even with truncation: "auto" explicitly configured, this provider now warns (once per response, via logger.warning) whenever the reported usage.input_tokens lands at or near the resolved context window for that model — a signal that the server likely truncated the prompt server-side. The warning names the model, the reported input_tokens, and the resolved context window, but never raises and never alters the response. It reuses the same per-model context-window resolution described above (server-reported max_model_len, clamped by an explicitly configured context_window), so it stays accurate on endpoints serving multiple models with different limits.

Local vs remote — single source of truth

base_url is the single source of truth for whether this provider instance is local or remote. The provider URL-parses it once at construction and caches the result; everything downstream (is_remote property, capability tagging in get_info() and list_models()) flows from that one decision.

  • Local — base_url resolves to localhost, 127.0.0.1, ::1, or 0.0.0.0. Capability tag: local. No auth required (api_key is ignored if it is the placeholder "EMPTY").
  • Remote — anything else (LAN IP, public hostname, RunPod / Modal / Anyscale / Lambda Labs URL, or a vLLM-backed proxy like OpenRouter/Together/Fireworks). Capability tag: remote. Bearer auth is attached when api_key is set.

The is_remote property is purely informational — it does not change how Bearer is attached (the OpenAI SDK does that whenever api_key is non-empty regardless of host). It exists so routing matrices and other downstream consumers can reason about the deployment shape.

Mixed local + remote (multi-instance)

To use both a local vLLM and a remote / hosted vLLM in the same session, configure two provider instances. Amplifier supports multiple named instances of the same provider module via the instance_id key:

# Default LOCAL instance — keeps the natural mount name "vllm"
[[providers]]
module = "amplifier-module-provider-vllm"
[providers.config]
base_url = "http://localhost:8000/v1"
default_model = "openai/gpt-oss-20b"

# Second instance — explicit `instance_id` makes it addressable as "vllm-remote"
[[providers]]
module = "amplifier-module-provider-vllm"
instance_id = "vllm-remote"
[providers.config]
base_url = "https://api.endpoints.anyscale.com/v1"  # or your hosted vLLM URL
api_key = "${VLLM_API_KEY}"
default_model = "meta-llama/Llama-3.3-70B-Instruct"

A routing matrix can then target each independently:

roles:
  reasoning:
    candidates:
      - provider: vllm-remote
        model: "meta-llama/Llama-3.3-70B-Instruct"
      - provider: vllm
        model: "openai/gpt-oss-20b"
  fast:
    candidates:
      - provider: vllm
        model: "openai/gpt-oss-20b"

The kernel validates that at most one entry per module omits instance_id (the "default" keeps the natural mount name); any additional entries must specify an instance_id. See amplifier-core/_session_init.py for the exact contract.

Usage Examples

Basic Chat

from amplifier_core import AmplifierSession

config = {
    "session": {
        "orchestrator": "loop-basic",
        "context": "context-simple"
    },
    "providers": [{
        "module": "provider-vllm",
        "config": {
            "base_url": "http://192.168.128.5:8000/v1",
            "default_model": "openai/gpt-oss-20b"
        }
    }]
}

async with AmplifierSession(config=config) as session:
    response = await session.execute("Explain quantum computing")
    print(response)

With Reasoning

config = {
    "providers": [{
        "module": "provider-vllm",
        "config": {
            "base_url": "http://192.168.128.5:8000/v1",
            "default_model": "openai/gpt-oss-20b",
            "reasoning": "high",  # Enable high-effort reasoning
            "reasoning_summary": "detailed"
        }
    }],
    # ... rest of config
}

async with AmplifierSession(config=config) as session:
    # Model will show internal reasoning before answering
    response = await session.execute("Solve this complex problem...")

With Tool Calling

config = {
    "providers": [{
        "module": "provider-vllm",
        "config": {
            "base_url": "http://192.168.128.5:8000/v1",
            "default_model": "openai/gpt-oss-20b"
        }
    }],
    "tools": [{
        "module": "tool-bash",  # Enable bash tool
        "config": {}
    }],
    # ... rest of config
}

async with AmplifierSession(config=config) as session:
    # Model can call tools autonomously
    response = await session.execute("List the files in the current directory")

Architecture

This provider uses the OpenAI SDK with a custom base_url pointing to your vLLM server. Since vLLM implements the OpenAI-compatible Responses API, the integration is clean and direct.

Key components:

  • VLLMProvider: Main provider class (handles Responses API calls)
  • _constants.py: Configuration defaults and metadata keys
  • _response_handling.py: Response parsing and content block conversion

Response flow:

ChatRequest → VLLMProvider.complete() → AsyncOpenAI.responses.create() →
→ vLLM Server → Response → Content blocks (Thinking + Text + ToolCall) → ChatResponse

Responses API Details

The vLLM provider uses the Responses API (/v1/responses) which provides:

  1. Structured reasoning: Separate reasoning blocks from response text
  2. Tool calling: Native function calling support
  3. Conversation state: Built-in multi-turn conversation handling
  4. Automatic continuation: Handles incomplete responses transparently

Tool format (vLLM Responses API):

{
  "type": "function",
  "name": "tool_name",
  "description": "Tool description",
  "parameters": {"type": "object", "properties": {...}}
}

Response structure:

{
  "output": [
    {"type": "reasoning", "content": [{"type": "reasoning_text", "text": "..."}]},
    {"type": "function_call", "name": "tool_name", "arguments": "{...}"},
    {"type": "message", "content": [{"type": "output_text", "text": "..."}]}
  ]
}

Debugging

Enable raw to attach the exact request params to the llm:request event:

config:
  raw: true    # Attach exact request params sent to the server

Check logs:

# Find recent session
ls -lt ~/.amplifier/projects/*/sessions/*/events.jsonl | head -1

# View raw requests
grep '"event":"llm:request:raw"' <log-file> | python3 -m json.tool

# View raw responses
grep '"event":"llm:response:raw"' <log-file> | python3 -m json.tool

Troubleshooting

Connection refused

Problem: Cannot connect to vLLM server

Solution:

# Check vLLM service status
sudo systemctl status vllm

# Verify server is listening
curl http://192.168.128.5:8000/health

# Check logs
sudo journalctl -u vllm -n 50

Tool calling not working

Problem: Model responds with text instead of calling tools

Verification:

  • ✅ vLLM version ≥0.10.1
  • ✅ Using Responses API (not Chat Completions)
  • ✅ Tools defined in request

Note: Tool calling works via Responses API without special vLLM flags. If it's not working, check the model supports tool calling.

No reasoning blocks

Problem: Responses don't include reasoning/thinking

Check:

  • Is reasoning parameter set in config? (minimal|low|medium|high)
  • Is the model a reasoning model? (gpt-oss supports reasoning)
  • Check raw debug logs to see if reasoning is in API response

Token usage shows zeros

For GPT-OSS models: Token accounting is automatic but requires vocab files.

How it works:

  • First use: Automatically downloads vocab files to ~/.amplifier/cache/vocab/
  • Subsequent uses: Uses cached files
  • No manual setup needed if you have internet access

What's computed:

  • Input tokens: Accurate count using Harmony's tokenization (matches model training format)
  • Output tokens: Approximate count based on visible output text
  • Limitation: Output count doesn't include hidden reasoning channels (REST API limitation)

If auto-download fails (offline/air-gapped):

# Manual setup for offline environments
mkdir -p ~/.amplifier/cache/vocab

# Download vocab files (on a machine with internet)
curl -sS -o ~/.amplifier/cache/vocab/o200k_base.tiktoken \
  https://openaipublic.blob.core.windows.net/encodings/o200k_base.tiktoken

curl -sS -o ~/.amplifier/cache/vocab/cl100k_base.tiktoken \
  https://openaipublic.blob.core.windows.net/encodings/cl100k_base.tiktoken

# Transfer ~/.amplifier/cache/vocab/ directory to offline machine
# Then set environment variable:
export TIKTOKEN_ENCODINGS_BASE=~/.amplifier/cache/vocab

Check logs for:

  • [TOKEN_ACCOUNTING] Downloading Harmony vocab files to ~/.amplifier/cache/vocab/... (first use)
  • [TOKEN_ACCOUNTING] Loaded Harmony GPT-OSS encoder (success)
  • [TOKEN_ACCOUNTING] Injected usage: input=X, output=Y (active)

Development

# Clone and install
git clone https://github.com/microsoft/amplifier-module-provider-vllm
cd amplifier-module-provider-vllm
uv pip install -e .

# Run tests
pytest tests/

# Check types and lint
make check

Testing

Run the offline suite with uv run pytest -q -m "not live". CI covers OpenAI SDK 2.8.1, 2.9.0, 2.53.0, and 3.22.1 on Python 3.11 and 3.12. The default lock targets the latest qualified SDK; retained 2.x environments remain covered. Real SDK JSON and SSE parsing tests use an in-memory HTTP transport, not paid model calls. The live model-list check remains explicitly deselected until a reachable vLLM endpoint is supplied.

Token accounting constructs validated SDK usage records with both cached-read and cache-write counts. Newer SDKs require the cache-write field even when vLLM does not report it. If a future usage-schema change prevents construction, the completed answer and original usage are retained with an explicit warning that corrected accounting is unavailable; a successful answer is not discarded. Error fixtures use the installed SDK's HTTP package (httpx on 2.x, httpx2 on 3.x), including when both are installed. Do not restore the old <2.9 cap to work around test fixtures: it conflicts with providers that require native Responses compaction. A future major SDK requires requalification. The supported minimum is the lowest qualified SDK, 2.8.1. CI names each SDK case and checks its effective version before tests. Test imports explicitly include the checkout root rather than depending on an editable installation. Responses tool invocation IDs are kept distinct from output-item IDs, including the second SDK request carrying a tool result. This aligns correlation with the Responses contract; it does not establish how every live vLLM server validates IDs.

See ai_working/vllm-investigation/ for comprehensive test scripts:

  • test_provider_simple.py - Basic provider functionality test
  • 06_test_responses_correct_format.py - Responses API format validation
  • 04_test_tool_calling.py - Tool calling verification

License

MIT

Contributing

Note

This project is not currently accepting external contributions, but we're actively working toward opening this up. We value community input and look forward to collaborating in the future. For now, feel free to fork and experiment!

Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit Contributor License Agreements.

When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.

This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.

Trademarks

This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.

Bounded output without a completion deadline

auto_continue is an optional settings-only provider configuration key, not a setup prompt. It defaults to true when omitted, preserving normal continuation of truncated responses. To disable continuation, set it explicitly in settings:

config:
  auto_continue: false

For a single bounded output operation, pass request_options={"auto_continue": False} to complete(). The per-call option takes precedence and never changes the mounted provider. The option is consumed locally and is not sent to the API. An incomplete response retains its partial content and usage, reports finish_reason="length", and is not retried with a larger output budget. Consumers must not treat that partial response as a complete summary. The provider advertises this optional contract as completion:auto_continue:v1. This option does not impose a time limit.

About

Reference implementation of a vLLM provider module for the Amplifier project

Resources

Code of conduct

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages