ClimateClaw is a Python service for building AI-assisted climate-data workflows. It provides the API, conversation handling, model prompting, persistent thread storage, and tool orchestration needed to support interactive work with climate data.
The project integrates LiteLLM-native prompting, MongoDB-backed conversation state, and MCP-based tool execution for retrieval, code execution, and domain-specific automation.
- FastAPI app with strict auth parity to the production Rust service (
/api/chatbot/*) - Streaming responses via LiteLLM/OpenAI-compatible SSE (
application/x-ndjson) with code + image variants - Persistent conversation threads in MongoDB and JSONL files (
threads/), plus per-user scratch space (cache/) - MCP manager that wires the backend to dedicated tool servers
- Docker compose stack that includes LiteLLM, the backend, and both MCP servers
- Comprehensive pytest suite covering auth, prompting, storage, litellm client helpers, and route matrices
- Web-search MCP server for ICON model + DKRZ/HPC docs with cancellable OpenAI Web Search calls
podmanordocker- Credentials & headers for the Freva auth services
Create .env (used by FastAPI, Docker, and MCP servers). See .env.example for guidance.
./prod.sh up -d --buildServices that start:
climateclaw: FastAPI app (debugpy toggle viaDEBUG=truefor remote debugging session)code-server: MCP server running the sandboxed Jupyter kernel and exposingcode_interpreterweb-search-server: MCP server doing web search via OpenAI API and exposingweb_searchrag-server: MCP server exposingget_context_from_resourceslitellm: LiteLLM proxy that readslitellm_config.yaml
Bind mounts expose /work, logs, threads, and shared cache to other Freva services.
The Docker Compose stack delegates local-model inference to an OpenAI-compatible vLLM endpoint. To run that endpoint on a GPU host with Apptainer, copy and configure the example environment file:
cp apptainer/.env.example apptainer/.envSet APPTAINER_BASE_DIR, VLLM_OCI_IMAGE, VLLM_MODEL, GPU and serving
settings, and VLLM_API_KEY in apptainer/.env. Then pull the configured
image and start vLLM:
./apptainer/vllm/run.sh pull
./apptainer/vllm/run.sh upUse ./apptainer/vllm/run.sh status, logs, and down to manage the service.
The runner uses --nv, persists Apptainer, Hugging Face, and vLLM caches below
APPTAINER_BASE_DIR, and supports additional vLLM options through
EXTRA_VLLM_ARGS.
Configure the main .env so LiteLLM can reach the service:
CLIMATECLAW_VLLM_MODEL_ID="<model-id>"
CLIMATECLAW_VLLM_API_BASE="http://<vllm-host>:8000/v1"
CLIMATECLAW_VLLM_API_KEY="<vllm-api-key>"<vllm-host> must be reachable from the LiteLLM container. This separation
also lets future deployments replace the locally managed Apptainer service with
dedicated LLM-inference infrastructure without changing ClimateClaw's client
interface.
podmanordocker
Create .env (used by FastAPI, Docker, and MCP servers). See .env.example for guidance.
./dev.sh up -d --build| Path | Purpose |
|---|---|
src/climateclaw/app.py |
FastAPI entrypoint, CORS policy, router registration, app lifespan hooks |
src/climateclaw/api/chatbot/* |
HTTP handlers for chat operations (availablechatbots, streamresponse, getthread, etc.) |
src/climateclaw/services/streaming/ |
LiteLLM client, orchestrator, stream variant definitions, heartbeat helpers |
src/climateclaw/services/storage/ |
MongoDB + disk-backed persistence (threads/ JSONL, cache/ scratch space) |
src/climateclaw/services/mcp/ |
MCP manager and MCP client |
src/climateclaw/services/authentication/ |
Authentication: DEV mode auth surpassing OIDC requirements |
src/climateclaw/core/ |
Settings, prompt assembly, logging, startup checks, available-model parsing |
src/climateclaw/tools/ |
MCP servers, auth helpers, header gate middleware |
prompt_library/ |
Baseline system prompts, summary prompts, and few-shot examples (JSONL) |
resources/ |
Documentation corpora used by the RAG tool (stableclimgen seed content) |
docker/ |
Dockerfiles for base, climateclaw and MCP servers |
scripts/ |
Dev utilities (dev_chat.py, dev_script.py, check_kernel_env.py) |
tests/ |
Pytest suite covering auth, prompting, streaming, storage, and endpoints |
litellm_config.yaml |
Source of truth for model catalog (consumed by available_chatbots()) |
Generated artifacts that persist across runs:
threads/(JSONL transcript per thread id)cache/{user_id}/{thread_id}(LLM-created files, plots, etc.)logs/(when mounted in Docker)
- FastAPI layer enforces auth via
AuthRequired(Bearer tokens validated againstx-freva-rest-url), derives stable UUIDv5 pseudonymous user IDs from usernames, and validates per-request headers. - LiteLLM proxy (
CLIMATECLAW_LITE_LLM_ADDRESS) provides OpenAI-compatible chat + embeddings endpoints; completions stream intoStreamVariantclasses that normalize assistant text, code blocks, tool hints, images, and server hints. - Persistence uses MongoDB for storing threads and user feedback.
- MCP Manager (
src/climateclaw/services/mcp/mcp_manager.py) connects to tool servers listed inCLIMATECLAW_AVAILABLE_MCP_SERVERS, discovers tools, exposes OpenAI function schemas to LiteLLM, and routes tool invocations with per-thread session ids. - MCP servers run as separate ASGI apps (dockerized). Requests flow through
header_gateso required headers (mongodb-uri,working-dir) become ContextVars before code executes. - Prompting loads baseline templates + few-shot examples per model and replays thread history (minus prompts, meta) to LiteLLM, matching the Rust semantics.
| Method | Path | Description | Notes |
|---|---|---|---|
GET |
/api/chatbot/ping |
Static ping stub | Placeholder |
GET |
/api/chatbot/docs |
Docs payload stub | Placeholder |
GET |
/api/chatbot/help |
Help payload stub | Placeholder |
GET |
/api/chatbot/availablechatbots |
Returns model names from litellm_config.yaml |
Requires auth |
GET |
/api/chatbot/newthread |
Generates a fresh thread_id |
Requires auth |
POST |
/api/chatbot/getthread |
Fetches thread contents omitting prompts + redundant StreamEnd variants | Requires auth |
POST |
/api/chatbot/getuserthreads |
Returns recent threads for authenticated user | JSON body: num_threads, page |
POST |
/api/chatbot/streamresponse |
Starts an SSE stream of StreamVariant JSON payloads |
Query params: thread_id, input (required), chatbot |
POST |
/api/chatbot/stop |
Initiates stopping of an active conversation | JSON body: thread_id; requires auth |
- Response type:
application/x-ndjson - Each
data:line is a JSON object withvariantdiscriminators (Assistant,Code,CodeOutput,CodeError,Image,ServerHint,StreamEnd, etc.). - Code tool calls stream incremental chunks while LiteLLM emits
tool_calls. When the MCP tool resolves, results are converted back into JSON events and appended to Mongo/disk storage. - The first chunk is a
ServerHintcarrying thethread_id; conversation variants are stored in-memory during streaming and flushed to MongoDB at the end, ensuring replay safety. - Clients can call
/api/chatbot/stop?thread_id=...to move a conversation intoSTOPPING; the streaming loop exits and cancels in-flight MCP requests (code, rag, web-search) via the sharedActiveRequestregistry.
- MongoDB (
mongodb_storage.py): canonical record for threads. Each document stores a UUIDv5 pseudonymoususer_id,thread_id, ISO timestamp, topic (summarized via LiteLLM), and serializedStreamVariantlist. cache/scratch:create_dir_at_cache()ensures each user/thread has a writable directory for generated files (plots, CSVs). Entries are sanitized if user IDs contain unsupported characters.- Prompt library:
prompt_library/baselinecontainsstarting_prompt.txt,summary_prompt.txt, andexamples.jsonl. GPT-5 models currently fall back to baseline prompts (warning logged). Customize by adding new prompt sets and updating_resolve_baseline_dir()/_resolve_gpt5_dir_or_placeholder(). - Resources:
resources/stableclimgenseeds the RAG MCP server. Drop additional corpora per library folder and list them inCLIMATECLAW_AVAILABLE_LIBRARIESinsidesrc/climateclaw/tools/rag/server.py.
- Code interpreter (
src/climateclaw/tools/code/server.py): spins up per-session Jupyter kernels, sanitizes input, enforces configurable timeouts, and injects Freva config via environment variables. Outputs include stdout/stderr, display data, and structured errors. - Web search server (
src/climateclaw/tools/web_search/server.py): calls OpenAI Web Search (gpt-4.1) constrained to ICON model + DKRZ/HPC docs. Honors request cancellation. - RAG server (
src/climateclaw/tools/rag/server.py): indexes documentation with custom loaders + splitters, stores embeddings in MongoDB (embeddings), and surfaces a single toolget_context_from_resources. LiteLLM requests embed queries through the same proxy (CLIMATECLAW_LITE_LLM_ADDRESS). - Header gate (
src/climateclaw/tools/header_gate.py): wraps each MCP ASGI app so critical headers become ContextVars and requests fail fast when missing/invalid (e.g., missing Mongo URI yields SSE-friendly JSON-RPC errors). - Manager (
src/climateclaw/services/mcp/mcp_manager.py): caches clients, discovers tool schemas, exports OpenAI function definitions, and pins MCP session ids to thread ids for deterministic tool contexts.
- Spin up dev stack:
./dev.sh up -d --build(FastAPI, rag, code, web-search, LiteLLM). Use./dev.sh up --buildto tail the app. Local Ollama models require an Ollama server running on the host at port11434(for example,ollama serve). - Unit/functional tests:
uv run pytestor focus, e.g.uv run pytest tests/test_auth.py -k bearer. - Integration: code interpreter:
CLIMATECLAW_CODE_SERVER_URL=http://localhost:8051 uv run pytest tests/full_integration_tests/test_code_interpreter.py -m integration. - Integration: web-search:
CLIMATECLAW_WEB_SEARCH_SERVER_URL=http://localhost:8052 uv run pytest tests/full_integration_tests/test_web_search.py -m integration. - Interactive chat:
uv run python scripts/dev_chat.pystarts a REPL that exercises the same orchestrator logic, persisting outputs to disk and optionally pointing at local MCP servers.
- Prod scaling:
./prod.sh up -d --buildgeneratesdocker-compose.scaled.yml+haproxy.cfgviagen_compose.py, then launches HAProxy in front of replicas. - Dev scaling:
./dev.sh --scale up -dproducesdocker-compose.dev.scaled.ymland matching HAProxy config. - Replica knobs: set
CLIMATECLAW_BACKEND_REPLICAS,CLIMATECLAW_LITELLM_REPLICAS, andCLIMATECLAW_{RAG|CODE|WEB_SEARCH}_REPLICAS(default 1). Only MCP servers listed inCLIMATECLAW_AVAILABLE_MCP_SERVERSare scaled. - Sticky routing: HAProxy pins backend and MCP tool traffic by
hdr(X-Freva-Thread-Id); LiteLLM stays round-robin. - Ports: HAProxy binds
CLIMATECLAW_TARGET_PORTfor the backend and4000for litellm; MCP frontends bind their configured ports (e.g., 8050/8051/8052) while container instances stay internal.
- Auth failures: verify headers include both
Authorizationandx-freva-rest-url. Inspect FastAPI logs for the exact HTTP status. - Missing models: ensure
litellm_config.yamlis readable and containsmodel_namekeys.available_chatbots()aborts the process if it cannot find any entries. - MCP issues: backend logs warn but continue when tool discovery fails; LiteLLM will simply not emit tool calls. Use
settings.AVAILABLE_MCP_SERVERSto enable/disable targets explicitly. - File access: Make sure
/workis mounted read-only where expected.
Copyright (C) 2025, freva-org
This project is free software: you can redistribute it and/or modify it under the terms of the GNU Affero General Public License as published by the Free Software Foundation, version 3 of the License.
See the LICENSE file for details.