Detect when an AI agent's behavior changes, even when your monitoring says everything is fine.
Traditional observability tells you what happened in this run. Agent Drift tells you how your agent is behaving differently from before.
BEFORE (prompt v17) AFTER (prompt v18)
Trajectory 2.3 steps 8.1 steps +245%
Tool calls 1.3 5.2 +286%
Latency 45ms 143ms +216%
Cost $0.0128 $0.0446 +249%
Blocked actions 0.0% 13.3% +13.3 pp
Eval score 0.91 0.86 -5%
WHAT CHANGED? prompt_version 17 → 18
Drift began at the same time prompt_version 18 first appeared.
The agent still succeeds, so no errors fire and no health checks fail. But it now loops through search → LLM three times, costs 3× more, and tries to send emails it never used to. Agent Drift finds this change on its own and shows you exactly what changed.
Status: early MVP (v0.1). The core loop works end to end: instrument → fingerprint → baseline → detect → explain → compare. APIs may change. Feedback and contributions are welcome.
- How it works
- Quickstart (2 minutes)
- Instrumenting your agent
- Guardrails
- The dashboard
- How drift is detected
- HTTP API
- Configuration
- Running with Docker and PostgreSQL
- Project layout
- Development
- Roadmap
- Contributing
- License
Your agent ──@observe / @tool / @llm / @guard──▶ Python SDK
│ (background batched HTTP)
▼
FastAPI ingest ──▶ SQLite / PostgreSQL
│
┌───────────────────────────┴───────────────────────────┐
▼ ▼
Behavioral fingerprint Drift engine (on read)
scalar · categorical · sequence · outcome baseline window vs current window
numeric · rate · distribution · sequence
│
▼
Explanation (onset + version changes)
│
▼
Dashboard & JSON API
- Observe. The SDK records each agent execution as a run made of spans (LLM calls, tool calls, retrievals, guardrail checks), plus version metadata (prompt version, model, git commit…).
- Represent. At ingest time, each run is converted into a behavioral fingerprint:
- Scalar: trajectory length, LLM calls, tool calls, retries, errors, latency, tokens, cost
- Categorical: tools used (with counts), models used
- Sequence: the ordered trajectory, e.g.
web_search → llm → web_search → llm - Outcome: success, evaluation score, guardrail blocks
- Baseline. An agent's normal behavior is its recent history: by default, the 200 runs before the latest 50.
- Detect. The latest window is compared against the baseline with plain, debuggable statistics (no ML).
- Explain. A change-point search estimates when the drift began, and version fields that changed around that time are listed as correlated changes. They are never presented as proven causes.
- Investigate. Open the dashboard, compare two versions side by side, and drill into individual runs.
Requires Python 3.10+.
git clone https://github.com/rajarshidattapy/drift.git
cd drift
# 1. Install the server and the SDK
pip install -e ./server -e ./sdk
# 2. Start the server (SQLite by default, stored in ./agentdrift.db)
agentdrift-serverIn a second terminal, run the included simulated research agent. It needs no API keys: the LLM and tools are fakes.
# Build a baseline with prompt v17
python examples/research_agent/agent.py --prompt-version 17 --runs 200
# Ship a "harmless" prompt tweak: v18 asks the agent to double-check every claim
python examples/research_agent/agent.py --prompt-version 18 --runs 60Open http://localhost:8000. research-agent is flagged with Drift detected. Click it to see what changed, then open Compare to diff v17 against v18.
pip install -e ./sdk # zero dependencies (PyPI package coming soon)import agentdrift
from agentdrift import observe, tool, llm, record_usage, score
agentdrift.configure(
endpoint="http://localhost:8000",
versions={"git_commit": "a91bc2f", "environment": "prod"}, # attached to every run
)
@tool # or @tool(name="web_search"), @tool(type="retrieval")
def web_search(query: str) -> list[str]:
...
@llm(model="claude-sonnet-5")
def summarize(text: str):
response = client.messages.create(...)
return response # token usage is read from `response.usage` automatically
@observe(agent="research-agent", prompt_version="18", agent_version="1.4.0")
def research(topic: str) -> str:
answer = summarize(web_search(topic))
score(0.91) # optional: attach an evaluation score (0..1)
return answerThat's all you need. Every call to research() becomes one run.
| API | What it does |
|---|---|
@observe(agent=..., **versions) |
Records each call as a run. Keyword arguments are version metadata. Works with sync and async functions. |
@tool, @llm |
Records calls as spans with input, output, timing, errors, and (for @llm) tokens, model, and cost. |
with agentdrift.run(agent, **versions): |
Context-manager form of @observe. |
with agentdrift.span(type, name) as s: |
Records any block as a span. type is one of llm, tool, retrieval, memory, guardrail, custom. |
record_usage(input_tokens, output_tokens, cost_usd, model) |
Adds token usage and cost to the innermost active span. |
score(value) |
Attaches an evaluation score to the current run. |
set_metadata(**kv) |
Attaches arbitrary metadata to the current run. |
flush() |
Blocks until queued traces are sent. It also runs automatically at exit. |
Recognized version fields: model, agent_version, prompt_version, tool_version, retriever_version, git_commit, environment. These are stored as first-class columns, so you can compare by any of them. You can also set them with environment variables such as AGENTDRIFT_PROMPT_VERSION or AGENTDRIFT_GIT_COMMIT. If model isn't set, it's inferred from the LLM spans.
Telemetry never breaks your agent. Traces are exported in batches from a background thread. If the server is unreachable, the SDK logs one warning and your agent keeps running.
Guardrails run a policy check before a tool executes. Every decision is recorded as a guardrail span, so a rise in blocked actions shows up as guardrail drift.
from agentdrift import Policy, guard, GuardrailBlocked
agentdrift.configure(
policy=Policy({
"delete_file": "block",
"send_email": {"action": "approval", "reason": "external side effect"},
"web_search": "allow",
}),
# Optional: return True to approve. Without a handler, "approval" means blocked.
approval_handler=lambda tool, args: ask_a_human(tool, args),
)
@guard("send_email")
@tool(name="send_email")
def send_email(to, body): ...
try:
send_email("someone@example.com", "hi")
except GuardrailBlocked as e:
print(e.reason)Policies can also be loaded from a file with Policy.from_file("policy.yaml") (.json works out of the box; .yaml needs pip install agentdrift[yaml]):
default: allow
tools:
delete_file:
action: block
send_email:
action: approval
reason: external side effectThe dashboard is served by the server at /. It's plain HTML and JavaScript with no build step.
- Agents: every agent with its status (Stable, Drift detected, Collecting baseline) and its top findings.
- Drift: the drift score, per-category severity ("why?"), findings, What changed? (estimated onset and correlated version changes), a before/after metric table, trajectory patterns, and tool mix.
- Compare: pick a version field (e.g.
prompt_version) and two values to get a behavioral diff between them. Think of it asgit difffor agent behavior. - Runs: recent runs with their trajectories. Click one to open a span timeline, where you can inspect each span's inputs, outputs, tokens, cost, and metadata.
- Versions: every version value seen, with run counts and first/last-seen times.
Every check is simple enough to reason about. The engine is ~600 lines of dependency-free Python in server/agentdrift_server/engine.
A metric is flagged only when the change is both large and statistically unlikely to be noise:
| Check | Metrics | Test | Flagged when |
|---|---|---|---|
| Numeric | trajectory length, LLM calls, tool calls, latency, tokens, cost, eval score | Welch z-score of the difference in means | |z| ≥ 3 and relative change ≥ threshold (e.g. 15% for steps, 20% for latency, 3% for eval score) |
| Rate | success rate, runs with errors, runs with retries, runs with blocked actions | Two-proportion z-test | |z| ≥ 3 and absolute change ≥ threshold (e.g. 2–3 pp) |
| Tool mix | share of calls per tool | Jensen–Shannon divergence | JSD ≥ 0.05 |
| Trajectory | step-transition bigrams (START → web_search, web_search → llm, …) |
Jensen–Shannon divergence | JSD ≥ 0.05 |
| Model | models used by LLM spans | Set difference | A new model appears in ≥ 10% of runs, or a model disappears |
Findings are grouped into categories (trajectory, tool, model, latency, error, cost, quality, guardrail), each graded low, medium, or high. The drift score (0–100) is a weighted summary: behavior 30%, quality 25%, performance 20%, cost 15%, safety 10%. It's a convenience for sorting, not the source of truth. The findings are.
Explanation. The engine runs a single change-point search (CUSUM-style) over the most-drifted signal to estimate the onset. It then compares the dominant value of every version field between the baseline and current windows. Changed fields are reported with timing, e.g. "Drift began 6 min after prompt_version 18 first appeared". If no version changed, it says so and points you to external causes: tool/API behavior, retrieval data, or inputs.
Minimum data: 20 baseline runs and 10 current runs. Below that, the agent shows Collecting baseline.
Interactive docs are at http://localhost:8000/docs.
| Method | Path | Description |
|---|---|---|
POST |
/api/v1/traces |
Ingest {"traces": [...]} (the SDK wire format). Idempotent on run_id. |
GET |
/api/v1/agents |
All agents with status, score, and top findings. |
GET |
/api/v1/agents/{name}/drift?window=50&baseline=200 |
Full rolling drift report with explanation. |
GET |
/api/v1/agents/{name}/compare?field=prompt_version&a=17&b=18 |
Behavioral diff between two versions. |
GET |
/api/v1/agents/{name}/runs?limit=50&offset=0 |
Recent runs. |
GET |
/api/v1/agents/{name}/versions |
Version values with run counts and first/last seen. |
GET |
/api/v1/runs/{run_id} |
A run with its spans and fingerprint. |
GET |
/api/v1/health |
Health check. |
You don't need the Python SDK. Any language can post traces in this format:
{
"traces": [{
"run_id": "run_82931",
"agent": "research-agent",
"started_at": "2026-09-16T09:42:00.000Z",
"ended_at": "2026-09-16T09:42:02.841Z",
"status": "success",
"eval_score": 0.91,
"versions": {"prompt_version": "18", "model": "claude-sonnet-5", "git_commit": "a91bc2f"},
"spans": [
{"span_id": "s1", "type": "tool", "name": "web_search", "status": "success",
"started_at": "2026-09-16T09:42:00.010Z", "ended_at": "2026-09-16T09:42:00.400Z"},
{"span_id": "s2", "type": "llm", "name": "summarize", "status": "success", "model": "claude-sonnet-5",
"input_tokens": 3821, "output_tokens": 911, "cost_usd": 0.031,
"started_at": "2026-09-16T09:42:00.410Z", "ended_at": "2026-09-16T09:42:02.800Z"}
]
}]
}Server
| Setting | Default | Description |
|---|---|---|
DATABASE_URL / --database-url |
sqlite:///./agentdrift.db |
Any SQLAlchemy URL. For PostgreSQL: postgresql+psycopg://user:pass@host/db (install with pip install -e "./server[postgres]"). |
AGENTDRIFT_HOST / --host |
127.0.0.1 |
Bind address. |
AGENTDRIFT_PORT / --port |
8000 |
Port. |
SDK (via agentdrift.configure(...) or environment variables)
| Option | Env var | Default | Description |
|---|---|---|---|
endpoint |
AGENTDRIFT_ENDPOINT |
http://localhost:8000 |
Server URL. |
enabled |
AGENTDRIFT_DISABLED=1 disables |
True |
Turns recording on or off. |
versions |
AGENTDRIFT_<FIELD> |
{} |
Version metadata merged into every run. |
capture_io |
– | True |
Records span/run inputs and outputs. Set to False for sensitive data. |
max_io_chars |
– | 2000 |
Truncation limit for captured inputs and outputs. |
policy, approval_handler |
– | None |
Guardrail configuration. |
exporter |
– | HTTP | A custom callable that receives each trace dict (useful for tests or other backends). |
⚠️ Security note: the MVP server has no authentication. By default it binds to localhost only. Don't expose it publicly, and considercapture_io=Falseif prompts or tool outputs contain sensitive data.
docker compose up --buildThis starts PostgreSQL and the server on http://localhost:8000. Point the SDK there as usual.
drift/
├── sdk/ # `agentdrift` Python SDK (no dependencies)
│ └── agentdrift/
│ ├── tracer.py # runs, spans, context propagation
│ ├── decorators.py # @observe, @tool, @llm
│ ├── guardrails.py # Policy, @guard, GuardrailBlocked
│ ├── client.py # background batched HTTP exporter
│ └── config.py
├── server/ # `agentdrift-server` (FastAPI + SQLAlchemy)
│ ├── agentdrift_server/
│ │ ├── engine/ # pure-Python drift engine (no I/O)
│ │ │ ├── fingerprint.py # trace → behavioral fingerprint
│ │ │ ├── detector.py # baseline vs current, severity, score
│ │ │ ├── explain.py # onset estimation + correlated version changes
│ │ │ └── stats.py # z-tests, JSD, change point
│ │ ├── app.py # HTTP API + dashboard hosting
│ │ ├── services.py # ingestion and queries
│ │ ├── db.py # agents, runs, spans tables
│ │ └── static/ # dashboard (HTML/CSS/JS, no build step)
│ └── Dockerfile
├── examples/research_agent/ # simulated agent that drifts on a prompt change
├── tests/ # engine, SDK and API tests
├── docs/idea.md # product requirements / vision
└── docker-compose.yml
pip install -e "./server[dev]" -e ./sdk
python -m pytest testsThe engine is pure functions over dicts. To experiment with a new detector, add it to engine/detector.py and cover it in tests/test_engine.py. There's a helper there that generates synthetic runs, and a test that confirms stable behavior is not flagged. Keep that test passing: false positives are what kill trust in a drift detector.
This MVP deliberately keeps the architecture boring. Next up, roughly in order (see docs/idea.md for the full vision):
- Asynchronous ingestion (Redis queue + worker) and persisted baselines and drift events
- Alerts: dashboard notifications, then webhooks and Slack
- LLM-as-judge evaluations
- OpenTelemetry export and ingestion (GenAI semantic conventions)
- Deployment/change events (e.g. "prompt v18 deployed at 09:39") alongside first-seen times
- Richer sequence drift: edit distance, loop detection, trajectory clustering
- CI mode: replay traces for a PR and fail on behavioral regressions
- Framework integrations (Claude Agent SDK, LangGraph, OpenAI Agents SDK, …)
- Authentication, projects, and API keys
- Kubernetes manifests; Prometheus metrics for Agent Drift itself
Out of scope: agent frameworks, agent builders, model hosting, vector databases, memory frameworks, prompt marketplaces, and generic APM.
Issues and pull requests are welcome. Good first contributions:
- A new example agent (coding agent, browser agent, customer-support agent)
- An integration that auto-instruments a popular agent framework or LLM client
- A new drift check, with a test showing it catches real drift and stays quiet on stable data
- Dashboard improvements
Please run python -m pytest tests before opening a PR, and keep changes focused.
