OpenAI-compatible proxy for AcademicAI (die KI-Initiative für Österreichs Universitäten im Rahmen des ACOmarket-Portfolios).
It exposes AcademicAI models on a local OpenAI-style API (default: http://127.0.0.1:11435/v1).
- Chat completions: ✅
- OpenAI Responses API (
POST /v1/responsesfor Codex CLI & Desktop): ✅ - Model list endpoint: ✅
- Health endpoint: ✅
- Cost status endpoint: ✅
- Local Request Cost Calculation: ✅ (Autonomous Decimal cost accounting, pricing cache & aggregation)
- Tool-call emulation (JSON-mode with TypeScript signatures & JSON repair): ✅
- SSE-style streaming emulation: ✅
- Daily Log Rotation (30 days retention): ✅
- E2E Test Port Isolation (runs on port 11436): ✅
- Automatic Prompt Caching Compatibility (Azure prefix caching): ✅
- Modular Domain Architecture & Modern ASGI Lifespan (
academicai.app): ✅ - Tiered Context Pricing & 128k Boundary Protection: ✅ (Structured ascending tiers, baseline rate guarantee, 128k client protection)
- OpenCode & OpenChamber Web/Mobile Harness (Tailscale): ✅ (see docs/opencode-openchamber.md)
- Automatic Prefix Caching: ✅ Supported natively. The proxy is aligned to merge system instructions at the very beginning of the first user message, maximizing Azure OpenAI prefix cache hit rates.
-
Autonomous Local Cost Calculation: ✅ Fully operational. The proxy dynamically caches model pricing from
/api/v1/llm/models, calculates exact request costs viaDecimalarithmetic, injects standardized response headers, and maintains persistent local aggregations (today,this_month,all_time,by_model,by_client). -
Tiered Context Pricing & 128k Boundary Protection: ✅ Upstream pricing tiers (Google Vertex AI / Azure OpenAI ShortCo
$\le 128\text{k}$ vs LongCo$> 128\text{k}$ ) are structured into ascending tiers inModelCatalog, with baseline rates guaranteed and official sources documented. The client recommendationcontext: 128000secures 100% Tier 1 rates. -
AcademicAI Backend Cost Endpoint (
/api/v1/cost/): 🔴 (Returns403 Forbiddendue to tenant permissionsACCESS_API_MONITOR_CREDIT). Local tracking completely bypasses this limitation.
AcademicAI does not provide native OpenAI function-calling/tool-calling in the same way OpenAI-compatible clients expect. This proxy emulates the tool flow so orchestrators (e.g. OpenClaw) can still run tools reliably.
Tool-calling is simulated, not native. That means the model is guided via prompt + JSON parsing, not by a backend-level function-calling engine.
In practice this works well, but there are limits:
- behavior is probabilistic (occasionally the model may answer in JSON style instead of ideal natural text)
- extra guardrails are needed to avoid unnecessary repeated tool calls
- reliability is generally lower than true native tool-calling APIs
So: good for practical use, but not mathematically deterministic.
Users often observe that tool calling through this proxy feels surprisingly fast. This is driven by three specific architectural choices:
- Azure OpenAI KV-Prefix Caching:
The proxy merges system instructions and tool definitions at the very beginning of the first user message. Because the AcademicAI backend runs on Azure OpenAI, stable prefix tokens (system context + tool signatures) trigger automatic KV-cache hits. This reduces Time-To-First-Token (TTFT) from seconds to milliseconds on repeated turns. - Single-Pass Minimal Output (JSON Mode):
Instead of a two-pass router or conversational tool descriptions, the backend is invoked inresponse_format: {type: "json_object"}. Modern models produce minimal JSON without pleasantries ({"action": "tool_call", ...}), emitting only 25–40 tokens per call. - Zero Heavy Framework Overhead:
Direct asynchronous HTTP transport viahttpxand Python standard library parsing eliminates the latency overhead of heavy abstraction layers.
While reliable for everyday agent tasks, emulating function calling over a text-only backend comes with structural trade-offs:
- Schema Compression Trade-offs:
To prevent context window explosion when dozens of tools are registered, tool schemas are compressed into high-density TypeScript-style signatures (_compact_tool_def), featuring concise enum unions, typed arrays (string[],number[]), defaults, and shallow objects ({query, tags}). Highly complex, deeply nested JSON schemas requiring deep object trees may still experience reduced parameter precision compared to native OpenAI function calling engines. tool_choicePrompt Enforcement:
While the proxy enforcestool_choice: "required"and specific tool targets by strictly forbidding{"action": "respond"}in system prompts and reminders, enforcement happens at the prompt layer rather than the engine sampler level.- Multi-Turn Role Flattening (
role: "tool"):
The underlying backend only acceptsuserandassistantroles. Tool results are flattened into user turns with standardized<tool_result id="..." name="...">XML tags. In deep multi-step loops (5+ sequential tool executions), this conversational history can sometimes cause attention drift, which the post-tool guard mitigates. - Probabilistic vs. Deterministic Parsing & JSON Repair:
Unlike native APIs where tool arguments are constrained by grammar-based token samplers, the model generates raw JSON text. The proxy pairs a multi-tier fallback parser (direct parse → markdown codeblock extraction → bracket-depth counter) with automatic JSON-repair sanitization (trailing comma stripping, unescaped control character leniency, and path escape fixes).
GET /health— Service & backend health checkGET /internal/cost-status— Cached cost snapshotGET /internal/opencode-models— opencode-compatible model cost map (authenticated)GET /v1/models— Dynamic model discoveryPOST /v1/chat/completions— Standard OpenAI Chat Completions APIPOST /v1/responses— OpenAI Responses API for OpenAI Codex CLI & Desktop
The proxy dynamically discovers and validates available models from the AcademicAI backend via GET /v1/models. Typical models supported include:
| Family | Model IDs | Key Features |
|---|---|---|
| OpenAI | gpt-4o, gpt-4o-mini, gpt-5, gpt-5-mini, gpt-5-nano, gpt-5.2, gpt-5.5, o3 |
Emulated tool-calling, Azure KV prefix caching, reasoning parameters |
| Anthropic | claude-opus-4-6, claude-opus-4-8 |
1M token context window, deep reasoning |
gemini-3.5-flash, gemini-3.1-flash-lite, gemini-3.1-pro-preview, gemini-2.5-pro |
1M token context window, multimodal capabilities | |
| Perplexity | sonar-pro, sonar-reasoning-pro |
Built-in search and citations |
| Mistral | Mistral-Large-3 |
256k context window |
Da der AcademicAI-Endpunkt /api/v1/cost/ für Standard-API-Clients 403 Forbidden (ACCESS_API_MONITOR_CREDIT) zurückgibt, führt der Proxy eine autonome lokale Kostenberechnung pro Request durch.
AcademicAI liefert über den Endpunkt /api/v1/llm/models die Kosteninformationen pro Modell in der Datenstruktur costs.
-
Maßeinheit & Währung: Preise in
costs(fürinput_tokensundoutput_tokens) sind in EUR pro 1.000 Tokens (1k Tokens) angegeben (z. B.0.00275fürgpt-4oInput = 0,00275 € / 1k Tokens). -
Berechnungsformel:
$$\text{input_rate} = \frac{\text{cost}}{1000}, \quad \text{output_rate} = \frac{\text{cost}}{1000}$$ $$\text{request_cost} = (\text{prompt_tokens} \times \text{input_rate}) + (\text{completion_tokens} \times \text{output_rate}) + \text{per_request_cost}$$ -
Hochpräzise Arithmetik: Sämtliche Berechnungen erfolgen mit Pythons
Decimal-Modul, um Rundungsfehler bei Mikro-Cents vollständig zu vermeiden. -
Kontext-Staffelung (Tiered Models): Modelle mit mehreren Preisstufen (
gpt-5.5,gemini-2.5-pro,gemini-3.1-pro-preview) nutzen die Basisstufe als Standard und markieren die Berechnung mitX-AcademicAI-Cost-Estimated: true.
AcademicAI aggregiert Modelle über Google Cloud (Vertex AI) und Microsoft Azure (Azure OpenAI Service). Für Modelle mit großem Kontextfenster (1M+) wenden die Upstream-Provider eine zweistufige Abrechnung an:
-
Google Gemini Pro (
gemini-2.5-pro):-
Tarifschwelle: Getrennt bei exakt 128k Tokens Prompt-Länge (
128.000 Tokens). -
Basis-Stufe (
$\le 128\text{k}$ Tokens): Input € 1,25 / 1M Tokens (0.00125€ / 1k), Output € 10,00 / 1M Tokens (0.01€ / 1k). -
Long-Context-Stufe (
$> 128\text{k}$ Tokens): Input € 2,50 / 1M Tokens (0.0025€ / 1k, verdoppelt), Output € 15,00 / 1M Tokens (0.015€ / 1k, 1,5x). - Quellen:
-
Tarifschwelle: Getrennt bei exakt 128k Tokens Prompt-Länge (
-
OpenAI via Microsoft Azure (
gpt-5.5):-
Tarifschwelle & Definition: Azure unterscheidet zwischen Short Context (
ShortCo) und Long Context (LongCo). Laut Microsoft bezieht sich Context hierbei strikt auf die Anzahl der Input-Tokens des einzelnen Requests ("number of input tokens in an individual request"), nicht auf das Modell-Kontextfenster oder Output-Tokens. -
Schwelle: Anfragen bis 128k Input-Tokens (
128.000 Tokens) fallen unter Short Context; Anfragen mit mehr als 128k Input-Tokens unter Long Context. -
Short-Context-Stufe (
$\le 128\text{k}$ Tokens): Input € 5,50 / 1M Tokens (0.0055€ / 1k), Output € 33,00 / 1M Tokens (0.033€ / 1k). -
Long-Context-Stufe (
$> 128\text{k}$ Tokens): Input € 11,00 / 1M Tokens (0.011€ / 1k, exakt 2x), Output € 49,50 / 1M Tokens (0.0495€ / 1k, exakt 1,5x). -
Quellen:
- Microsoft Foundry / Azure OpenAI Model Concepts — Short context and long context
- Azure OpenAI Service Pricing
-
Azure Retail Prices API (Meters
5.5 ShortCo inp Dz/5.5 LongCo inp Dz).
-
Tarifschwelle & Definition: Azure unterscheidet zwischen Short Context (
Note
Praxis-Standard: 128k-Limit garantiert Basis-Tarif (Tier 1)
In unserer Client-Konfiguration für OpenCode/OpenChamber (siehe unten) setzen wir standardmäßig ein Cap von limit.context: 128000 (128k Tokens). Dies garantiert, dass alle Anfragen ausnahmslos in der günstigsten Basis-Preisstufe (Tier 1) abgerechnet werden und die teurere Long-Context-Stufe (Tier 2) niemals getriggert wird. Im ModelCatalog (data/model_catalog.json) sind dennoch sämtliche Tiers vollständig und sortiert hinterlegt.
Jede erfolgreiche Anfrage über /v1/chat/completions (sowohl non-streaming als auch streaming) sowie über /v1/responses liefert standardisierte Kosten- und Token-Header zurück:
| Header | Beschreibung | Beispiel |
|---|---|---|
X-AcademicAI-Request-Cost |
Gesamtkosten des Requests | 0.000825 |
X-AcademicAI-Input-Cost |
Berechnete Input-Token-Kosten | 0.00033 |
X-AcademicAI-Output-Cost |
Berechnete Output-Token-Kosten | 0.000495 |
X-AcademicAI-Prompt-Tokens |
Tatsächliche Prompt-Tokens | 120 |
X-AcademicAI-Completion-Tokens |
Tatsächliche Completion-Tokens | 45 |
X-AcademicAI-Cost-Currency |
Währung (Standard: EUR) | EUR |
X-AcademicAI-Cost-Estimated |
true, falls Modellpreise geschätzt/gestaffelt |
false |
Hinweis: Standard-OpenAI-Response-Payloads bleiben 100 % unverändert und frei von proprietären Feldern, um die Kompatibilität mit Clients wie Cursor, Codex oder OpenClaw zu garantieren.
Der Proxy erfasst alle verfügbaren Modelle und deren Preise einmalig vom Endpunkt /api/v1/llm/models und persistiert den Modellkatalog atomar als JSON-Datei in data/model_catalog.json mit einer Lebensdauer (TTL) von 24 Stunden (86.400 Sekunden):
- Zero-Latency Model Discovery (
GET /v1/models): Der Proxy liefert die Modell-Liste direkt aus dem im Speicher gehaltenen Katalog im Standard-OpenAI-Format (to_openai_models_response()). Dies eliminiert den 200–500ms langen Upstream-Roundtrip bei jedem Start von Agenten-Tools (OpenCode, Codex, OpenClaw) und bietet vollständige Offline-Resilienz gegen Upstream-Ausfälle. - Metadaten & Tokengrenzen: Speichert neben Preisen auch
context_window(z. B. bis zu 1.050.000 Tokens) undoutput_token_limit(max_tokens). - Startup ohne Latenz: Beim Start des Proxies wird der Katalog sofort aus der lokalen JSON-Datei geladen.
- Background Refresh: Nach Ablauf der 24 Stunden wird der Refresh asynchron im Hintergrund ausgelöst, ohne den Client-Request zu blockieren.
- Fehlertoleranz: Sollte die AcademicAI-Modell-API temporär nicht erreichbar sein, greift der Proxy transparent auf den zuletzt gespeicherten Stand zurück.
opencode berechnet Session-Kosten aus model.cost × tokens. Für den Custom-Provider academicai fehlen diese Preisdaten, sodass der $ Spent-Zähler sonst bei 0 bliebe. Dafür gibt es zwei Bausteine:
GET /internal/opencode-models(authentifiziert viaverify_key) liefert den lokalenModelCatalogin opencodes Konvention (Preis pro 1.000.000 Token):Für gestaffelte Modelle kommt zusätzlich{"currency": "EUR", "models": {"gpt-5-mini": {"input": 0.28, "output": 2.2}}}context_over_200khinzu.- opencode-Plugin
integrations/opencode/academicai-cost-plugin.jsruft diesen Endpoint beim Start auf und injiziertcostin die konfiguriertenacademicai-Modelle. Installation und Optionen:integrations/opencode/README.md. - opencode 2.x nutzt statt des Plugins den Proxy-Config-Sync: der Proxy schreibt beim Start die
cost-Felder direkt in die globale opencode-Config (Begründung und Mechanik: ADR 0002); das Plugin bleibt nur für 1.18.x-Bestandsinstallationen relevant.
Grenzen (bewusst, siehe ADR docs/decisions/0001-opencode-cost-integration.md und ADR 0002):
- Werte sind EUR, opencode rendert
$- bewusst 1:1 ohne Umrechnung. -
per_request_costund Cache-Preise (cache_read/cache_write) sind im opencode-Schema nicht abbildbar;$ Spentkann dadurch geringfügig unterschätzen. -
Long-Context-Tarif: opencode 2.0.22 wendet
context_over_200kkorrekt als Tier ab 200k Input-Tokens an (Kostenformel v2:(nicht-gecachter Input × input + Cache-Read × cache_read + (Output + Reasoning) × output) / 1e6); der Proxy meldet die Stufe ab 128k, die Client-Standard-limit.context-Konvention von 128k hält die Rechnung bei Baseline-Raten. opencode 1.18.31 ignoriert das Feld (historisch, hermetisch verifiziert). -
Monats-/Tagesaggregation (
/internal/cost-status):today/this_monthlaufen auf Wien-Lokalzeit (Reset implizit am 1. um 00:00 Ortszeit, DST-korrekt); Zeitstempel bleiben UTC. - Die Anzeige wirkt nur für neue Assistant-Messages; Alt-Sessions bleiben bei
0.
Der Proxy aggregiert die Kosten serverseitig in-memory und persistiert sie atomar in data/local_cost_cache.json.
Der Endpunkt GET /internal/cost-status liefert:
backend_cost_monitoring: Status der AcademicAI-Kostenüberwachung (falls vorhanden).local_cost_tracking:all_time: Gesamtkosten, Token-Summen und Request-Count.today: Aggregation für den aktuellen Wien-Lokaltag (0:00–24:00 Ortszeit, DST-korrekt).this_month: Aggregation für den aktuellen Wien-Lokalmonat (impliziter Reset am 1., 00:00 Ortszeit).by_model: Aufschlüsselung pro Modell-ID.by_client: Aufschlüsselung nach anonymisiertem Client-Hash (client_<sha256[:8]>).model_catalog: Status des Modellkatalogs (models_cached,last_refreshed_at,is_stale,ttl_seconds,catalog_file).recent_requests: Ringpuffer der letzten 500 Requests (streng datenschutzkonform: nur Metadaten, keine Prompts, Completions oder API-Keys!).
In der .env konfigurierbar:
| Variable | Typ | Default | Beschreibung |
|---|---|---|---|
ACADEMICAI_ENABLE_LOCAL_COST_TRACKING |
bool | true |
Aktiviert die lokale Kostenberechnung & Header |
ACADEMICAI_COST_CURRENCY |
str | EUR |
Währung für Abrechnung & Response-Header (Standard: EUR) |
ACADEMICAI_MODEL_CATALOG_FILE |
str | data/model_catalog.json |
Pfad zum persistenten Modellkatalog |
ACADEMICAI_MODEL_CATALOG_TTL_SECONDS |
int | 86400 |
Gültigkeitsdauer des Modellkatalogs in Sekunden (24 Stunden) |
ACADEMICAI_LOCAL_COST_CACHE_FILE |
str | data/local_cost_cache.json |
Pfad zur lokalen Aggregationsdatei |
ACADEMICAI_LOCAL_COST_HISTORY_LIMIT |
int | 500 |
Maximale Einträge im Ringpuffer der Request-Historie |
ACADEMICAI_OPENCODE_CONFIG_FILE |
str | (leer) | Opencode-Config-Pfad-Override für den v2 Config-Sync (Hierarchie: Override > XDG_CONFIG_HOME/opencode/opencode.json > ~/.config/opencode/opencode.json); Details: ADR 0002 |
Client -> Proxy (Bearer):
Authorization: Bearer <YOUR_PROXY_API_KEY>
Proxy -> AcademicAI backend:
X-Client-ID: <ACADEMICAI_CLIENT_ID>X-Client-Secret: <ACADEMICAI_CLIENT_SECRET>
Configure these values in .env (never commit real secrets).
Startup fails fast when ACADEMICAI_PROXY_API_KEY is missing, insecure, or too short.
Keep repository content generic. Put tenant-specific values outside the repository:
.envwith endpoint, client ID/secret, proxy API key
Templates are provided under docs/tenant-template/.
pip install -r requirements.txt
cp .env.example .env
# edit .envRun:
py server.py
# or
.\start_server.ps1
# optional controlled stop
.\stop_server.ps1Der Proxy unterstützt OpenAI Codex (CLI und Desktop) über den standardkonformen Endpunkt POST /v1/responses (wire_api = "responses").
Konfiguration in ~/.codex/config.toml bzw. %USERPROFILE%\.codex\config.toml:
model = "gpt-4o"
model_provider = "academicai"
[model_providers.academicai]
name = "AcademicAI"
base_url = "http://127.0.0.1:11435/v1"
wire_api = "responses"
env_key = "ACADEMICAI_PROXY_API_KEY"
supports_websockets = falseAusführliche Details zu Tools, Multi-Turn-Roundtrips und Modellwahl findest du in docs/codex.md.
Der Proxy dient als primäres LLM-Backend für OpenCode und das Web-/PWA-Frontend OpenChamber via @ai-sdk/openai-compatible.
- Konfigurationsanleitung: Vollständige Einrichtung und Tailscale-Sicherheitsarchitektur siehe
docs/opencode-openchamber.md. - Modell-Charakteristiken: In
opencode.jsonsolltenlimit.context,limit.output,tool_callundreasoningstets explizit hinterlegt werden, um konservative Fallbacks (4k/8k) zu vermeiden. - Empfehlung: 128k-Kontextgrenze: Da Upstream-Tarife (wie Azure OpenAI und Google Vertex) sowie AcademicAI oberhalb von 128k Tokens signifikant teurer werden (Long-Context-Stufe / Tier 2 mit doppeltem Input-Preis), wird für alle Modelle ein striktes Cap auf
context: 128000(128k Tokens) gesetzt. Dies hält alle Sessions garantiert im günstigsten Basistarif, schützt das Budget vor unerwarteten Kostensprüngen und bietet mit ~400–500 Buchseiten Text ausreichend Raum für produktive Coding-Workflows.
The repository now includes a separate local test setup:
.env.localtestfor local test defaults.env.localtest.exampleas template.\start_test_server.ps1to start the proxy with local test settings.\run_local_tests.ps1 -Mode offlinefor local/offline regression tests.\run_local_tests.ps1 -Mode e2efor end-to-end proxy tests against AcademicAI
Offline mode:
- does not require real AcademicAI backend credentials
- validates hardening, request validation, guard logic, and humanization helpers
E2E mode:
- requires
ACADEMICAI_BASE_URL,ACADEMICAI_CLIENT_ID, andACADEMICAI_CLIENT_SECRETin.env.localtest - uses
ACADEMICAI_TEST_PROXY_API_KEY/ACADEMICAI_TEST_BASE_URLfrom.env.localtestfor the local proxy side
Quick start for local testing:
.\run_local_tests.ps1 -Mode offlineThe repository includes a versatile CLI tool test_models_connectivity.py to inspect registered models, pricing, and connectivity:
# 1. Quick overview of available upstream models & pricing (no completion tokens used):
python test_models_connectivity.py --upstream --list
# 2. Test a single model directly against the AcademicAI upstream backend:
python test_models_connectivity.py --upstream -m gpt-5-mini
# 3. Test all models via the running local proxy (port 11435):
python test_models_connectivity.py
# 4. Test a specific model family via local proxy:
python test_models_connectivity.py -m claudeKey features:
- Dual-Mode: Test through local proxy or directly against upstream BOKU AcademicAI API (
--upstream/-u). - Precise Error Diagnostics: Formats OpenAI-style
error.message, BOKUmeta.error.message(e.g.Cost limit reached), and FastAPI details without masking. - Model Listing: Formats context window, output token limit, and normalized input/output costs in €/1M tokens (
--list/-l).
ACADEMICAI_PROXY_API_KEYis mandatory and must be changed from insecure placeholders.- API key values shorter than 16 chars are rejected at startup.
- Debug dumps are secret-redacted (
Authorization, tokens, client secrets). POST /v1/chat/completionsenforces request shape and size limits.- Per-minute rate limiting is enabled by default (
ACADEMICAI_RATE_LIMIT_PER_MINUTE=120). GET /healthincludes backend status and reportsdegradedif backend check fails.
This repository is structured so both humans and LLM agents can install it reliably. Recommended approach: give your coding agent the GitHub URL and ask it to perform setup + verification.
Suggested instruction you can paste to an agent:
Install and verify this repository as a local service:
1) clone repo
2) create .env from .env.example
3) fill ACADEMICAI_BASE_URL, ACADEMICAI_CLIENT_ID, ACADEMICAI_CLIENT_SECRET, ACADEMICAI_PROXY_API_KEY
4) install dependencies
5) start server
6) verify /health and /v1/models with Bearer auth
7) run pytest
8) report final status and exact local run command
Minimum verification checks:
GET /healthreturns{"status":"ok"...}GET /v1/modelsworks withAuthorization: Bearer <ACADEMICAI_PROXY_API_KEY>py -m pytest -qpasses
Required:
modelmessages
Forwarded when set:
streamtemperaturemax_tokensmax_completion_tokensfrequency_penaltypresence_penaltyreasoning_effortverbosityseedstopresponse_format(dict only; tool-mode forces{ "type": "json_object" })extra_body.tailoredAiId
Tool emulation input (not forwarded natively):
tools(alias:functions)tool_choice
Validation and protection behavior:
- Invalid JSON body:
400 - Invalid request shape (e.g. wrong
messagestype):422 - Oversized payload/tool schema/message content:
413 - Rate limit exceeded:
429
OpenAI Responses API endpoint designed for OpenAI Codex CLI and Desktop (wire_api = "responses"):
Required:
model(e.g.gpt-4o,gpt-4o-mini, etc.)input(list of structured input items or plain text string)
Supported input item types:
type: "message": Standard conversational turn. Roles supported:user,assistant,system,developer. Text content can be a plain string or array of parts (type: "input_text"ortype: "output_text").type: "function_call": Tool calls emitted in prior turns (call_id,name,arguments).type: "function_call_output": Output returned from client-side execution in Codex's local sandbox (call_id,output).
Optional:
instructions: Top-level system prompt instructions.tools: List of tool definitions. Supports both flat schema ({"type": "function", "name": "...", "parameters": {...}}) and nested OpenAI schema ({"type": "function", "function": {...}}).stream: Boolean (truefor SSE streaming wire events,falsefor non-streaming JSON).temperature: Temperature override.max_output_tokens: Maximum completion tokens to generate.
Token Accounting & Metadata:
- Emits
input_tokensandoutput_tokens(strictly required by OpenAI Codex CLI's Rust parser) as well asprompt_tokens,completion_tokens, andtotal_tokensinusage.
These defaults apply only when the client did not set the field explicitly.
ACADEMICAI_DEFAULT_CHAT_TEMPERATURE=0.6ACADEMICAI_DEFAULT_CHAT_VERBOSITY=medium(gpt-5* only)ACADEMICAI_DEFAULT_TOOL_TEMPERATURE=0.1ACADEMICAI_DEFAULT_TOOL_VERBOSITY=low(gpt-5* only)ACADEMICAI_DEFAULT_TOOL_REASONING_EFFORT=low(gpt-5* only)
Rule:
- Human conversation without tool mode -> chat defaults
- Tool mode (
tools/functionspresent) -> tool defaults - Explicit client fields always win
You can enable a second LLM pass that rewrites structured/tool-derived output into natural human text.
- Active only for human channels
- Active only in tool mode
- Skipped when the model emits an actual tool call (
finish_reason: tool_calls)
Env flags:
ACADEMICAI_ENABLE_HUMANIZATION_PASS=true|falseACADEMICAI_HUMANIZATION_MODEL=<optional override>(default: same model)ACADEMICAI_HUMANIZATION_TEMPERATURE=0.2
Cost monitoring is enabled by default (ACADEMICAI_ENABLE_COST_MONITORING=true). It performs lazy/background refreshes of the AcademicAI GET /api/v1/cost snapshot and adds monitoring response headers. Independent of this, the local per-request cost calculation (default on) always works and does not require the upstream permission below.
Env flags:
ACADEMICAI_ENABLE_COST_MONITORING=true|false(default:false)ACADEMICAI_COST_CACHE_FILE=./cost_cache.jsonACADEMICAI_COST_CACHE_TTL_SECONDS=600ACADEMICAI_COST_REFRESH_TIMEOUT_SECONDS=8
Behavior when enabled:
- Proxy lazily/background-refreshes cost cache from AcademicAI
GET /api/v1/cost. - Adds response headers on chat completions (
X-AcademicAI-Total-Cost,X-AcademicAI-Total-Clients,X-AcademicAI-Cost-Entries,X-AcademicAI-Cost-Updated-At,X-AcademicAI-Cost-Stale). - Exposes
GET /internal/cost-status.
Important prerequisites (per AcademicAI API docs):
- API client permission
ACCESS_API_MONITOR_CREDITis required for/api/v1/cost. - Without that permission, the endpoint returns
403and cache stays empty/stale.
If you want proxy defaults to control style, do not hard-set these in OpenClaw for this provider:
temperatureverbosityreasoning_effort
- Proxy reads
toolsfrom request - Injects tool schema/instructions into prompt
- Forces JSON response mode
- Parses model JSON into either:
- single tool call:
{"action":"tool_call",...} - multi-step tool calls:
{"action":"tool_calls","calls":[...]} - normal assistant text (
{"action":"respond",...})
- single tool call:
- Converts tool call(s) into OpenAI
tool_callsresponse (finish_reason: tool_calls) - Upstream orchestrator executes tool and sends
role=toolfollow-up
When tool mode is active and the latest message already has role=tool,
the proxy injects a short guard instruction that prefers a final user-facing
answer and discourages unnecessary additional tool calls.
This reduces accidental re-tooling loops while still allowing another tool call if the latest tool result is clearly incomplete.
curl -s http://127.0.0.1:11435/healthcurl -s -H "Authorization: Bearer <ACADEMICAI_PROXY_API_KEY>" \
http://127.0.0.1:11435/v1/modelscurl -s -H "Authorization: Bearer <ACADEMICAI_PROXY_API_KEY>" \
-H "Content-Type: application/json" \
http://127.0.0.1:11435/v1/chat/completions \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Hello!"}
]
}'curl -s -H "Authorization: Bearer <ACADEMICAI_PROXY_API_KEY>" \
-H "Content-Type: application/json" \
http://127.0.0.1:11435/v1/responses \
-d '{
"model": "gpt-4o",
"instructions": "You are a concise assistant.",
"input": [
{"type": "message", "role": "user", "content": "Say OK"}
]
}'Run the complete offline regression test suite (186 unit and integration tests):
.\run_local_tests.ps1 -Mode offlineOr via pytest directly:
pytest -qTo run end-to-end tests against the live AcademicAI backend (requires credentials configured in .env.localtest):
.\run_local_tests.ps1 -Mode e2eacademicai-proxy/
academicai/
__init__.py
app.py # FastAPI application factory & ASGI lifespan
auth.py # AcademicAI authentication & header injection
config.py # Typed settings & environment parsing
cost_calculation.py # ModelCatalog & Decimal request cost calculation
cost_monitoring.py # Atomic cache & cost status
errors.py # Standardized OpenAI error mapping
humanization.py # Target channel detection & 2nd-pass rewriting
local_cost_tracker.py # LocalCostStore, aggregations & ring-buffer history
logging_config.py # Rotating file handlers & uvicorn wiring
opencode_config_sync.py # Proxy-driven opencode config sync (v2 cost fields)
opencode_costs.py # opencode-compatible per-1M cost map projection
provider.py # HTTP transport to AcademicAI backend
request_guards.py # Inbound payload validation & rate limiting
responses.py # OpenAI Responses API normalization & SSE serialization
runtime.py # Process lifecycle & backend health checks
tool_emulation.py # TypeScript signatures, repair & post-guard
transformation.py # Message role normalization & text extraction SSOT
docs/
architecture/ # Concept & modularization architecture docs
archive/ # Legacy archive artifacts
codex.md # OpenAI Codex CLI & Desktop setup guide
decisions/ # Architecture Decision Records (ADRs)
opencode-openchamber.md # OpenCode & OpenChamber Web/Mobile (Tailscale) setup guide
system-map/ # ICM-aligned agent architecture map
tenant-template/ # Environment templates
integrations/
opencode/ # opencode cost plugin (1.18.x) & install docs
server.py # Slim CLI runner & compatibility layer
start_server.ps1 # Controlled service startup
stop_server.ps1 # Controlled service shutdown
run_local_tests.ps1 # Offline and E2E test runner
tests/ # Comprehensive offline test suite (270+ tests)
requirements.txt
README.md
If your agent or team uses this proxy in production and you patch a bug, please contribute it upstream so others benefit too:
(If the repository URL changes, update this section accordingly.)
MIT (see LICENSE).