Skip to content

Repository files navigation

Agent Drift

Detect when an AI agent's behavior changes, even when your monitoring says everything is fine.

image

Traditional observability tells you what happened in this run. Agent Drift tells you how your agent is behaving differently from before.

                BEFORE (prompt v17)     AFTER (prompt v18)
Trajectory      2.3 steps               8.1 steps          +245%
Tool calls      1.3                     5.2                +286%
Latency         45ms                    143ms              +216%
Cost            $0.0128                 $0.0446            +249%
Blocked actions 0.0%                    13.3%              +13.3 pp
Eval score      0.91                    0.86                 -5%

WHAT CHANGED?   prompt_version 17 → 18
                Drift began at the same time prompt_version 18 first appeared.

The agent still succeeds, so no errors fire and no health checks fail. But it now loops through search → LLM three times, costs 3× more, and tries to send emails it never used to. Agent Drift finds this change on its own and shows you exactly what changed.

Status: early MVP (v0.1). The core loop works end to end: instrument → fingerprint → baseline → detect → explain → compare. APIs may change. Feedback and contributions are welcome.


Contents


How it works

 Your agent ──@observe / @tool / @llm / @guard──▶  Python SDK
                                                      │  (background batched HTTP)
                                                      ▼
                                             FastAPI ingest ──▶ SQLite / PostgreSQL
                                                      │
                          ┌───────────────────────────┴───────────────────────────┐
                          ▼                                                       ▼
               Behavioral fingerprint                                   Drift engine (on read)
      scalar · categorical · sequence · outcome            baseline window vs current window
                                                           numeric · rate · distribution · sequence
                                                                        │
                                                                        ▼
                                                      Explanation (onset + version changes)
                                                                        │
                                                                        ▼
                                                           Dashboard & JSON API
  1. Observe. The SDK records each agent execution as a run made of spans (LLM calls, tool calls, retrievals, guardrail checks), plus version metadata (prompt version, model, git commit…).
  2. Represent. At ingest time, each run is converted into a behavioral fingerprint:
    • Scalar: trajectory length, LLM calls, tool calls, retries, errors, latency, tokens, cost
    • Categorical: tools used (with counts), models used
    • Sequence: the ordered trajectory, e.g. web_search → llm → web_search → llm
    • Outcome: success, evaluation score, guardrail blocks
  3. Baseline. An agent's normal behavior is its recent history: by default, the 200 runs before the latest 50.
  4. Detect. The latest window is compared against the baseline with plain, debuggable statistics (no ML).
  5. Explain. A change-point search estimates when the drift began, and version fields that changed around that time are listed as correlated changes. They are never presented as proven causes.
  6. Investigate. Open the dashboard, compare two versions side by side, and drill into individual runs.

Quickstart (2 minutes)

Requires Python 3.10+.

git clone https://github.com/rajarshidattapy/drift.git
cd drift

# 1. Install the server and the SDK
pip install -e ./server -e ./sdk

# 2. Start the server (SQLite by default, stored in ./agentdrift.db)
agentdrift-server

In a second terminal, run the included simulated research agent. It needs no API keys: the LLM and tools are fakes.

# Build a baseline with prompt v17
python examples/research_agent/agent.py --prompt-version 17 --runs 200

# Ship a "harmless" prompt tweak: v18 asks the agent to double-check every claim
python examples/research_agent/agent.py --prompt-version 18 --runs 60

Open http://localhost:8000. research-agent is flagged with Drift detected. Click it to see what changed, then open Compare to diff v17 against v18.

Instrumenting your agent

pip install -e ./sdk        # zero dependencies (PyPI package coming soon)
import agentdrift
from agentdrift import observe, tool, llm, record_usage, score

agentdrift.configure(
    endpoint="http://localhost:8000",
    versions={"git_commit": "a91bc2f", "environment": "prod"},  # attached to every run
)

@tool                                   # or @tool(name="web_search"), @tool(type="retrieval")
def web_search(query: str) -> list[str]:
    ...

@llm(model="claude-sonnet-5")
def summarize(text: str):
    response = client.messages.create(...)
    return response                     # token usage is read from `response.usage` automatically

@observe(agent="research-agent", prompt_version="18", agent_version="1.4.0")
def research(topic: str) -> str:
    answer = summarize(web_search(topic))
    score(0.91)                         # optional: attach an evaluation score (0..1)
    return answer

That's all you need. Every call to research() becomes one run.

API What it does
@observe(agent=..., **versions) Records each call as a run. Keyword arguments are version metadata. Works with sync and async functions.
@tool, @llm Records calls as spans with input, output, timing, errors, and (for @llm) tokens, model, and cost.
with agentdrift.run(agent, **versions): Context-manager form of @observe.
with agentdrift.span(type, name) as s: Records any block as a span. type is one of llm, tool, retrieval, memory, guardrail, custom.
record_usage(input_tokens, output_tokens, cost_usd, model) Adds token usage and cost to the innermost active span.
score(value) Attaches an evaluation score to the current run.
set_metadata(**kv) Attaches arbitrary metadata to the current run.
flush() Blocks until queued traces are sent. It also runs automatically at exit.

Recognized version fields: model, agent_version, prompt_version, tool_version, retriever_version, git_commit, environment. These are stored as first-class columns, so you can compare by any of them. You can also set them with environment variables such as AGENTDRIFT_PROMPT_VERSION or AGENTDRIFT_GIT_COMMIT. If model isn't set, it's inferred from the LLM spans.

Telemetry never breaks your agent. Traces are exported in batches from a background thread. If the server is unreachable, the SDK logs one warning and your agent keeps running.

Guardrails

Guardrails run a policy check before a tool executes. Every decision is recorded as a guardrail span, so a rise in blocked actions shows up as guardrail drift.

from agentdrift import Policy, guard, GuardrailBlocked

agentdrift.configure(
    policy=Policy({
        "delete_file": "block",
        "send_email": {"action": "approval", "reason": "external side effect"},
        "web_search": "allow",
    }),
    # Optional: return True to approve. Without a handler, "approval" means blocked.
    approval_handler=lambda tool, args: ask_a_human(tool, args),
)

@guard("send_email")
@tool(name="send_email")
def send_email(to, body): ...

try:
    send_email("someone@example.com", "hi")
except GuardrailBlocked as e:
    print(e.reason)

Policies can also be loaded from a file with Policy.from_file("policy.yaml") (.json works out of the box; .yaml needs pip install agentdrift[yaml]):

default: allow
tools:
  delete_file:
    action: block
  send_email:
    action: approval
    reason: external side effect

The dashboard

The dashboard is served by the server at /. It's plain HTML and JavaScript with no build step.

  • Agents: every agent with its status (Stable, Drift detected, Collecting baseline) and its top findings.
  • Drift: the drift score, per-category severity ("why?"), findings, What changed? (estimated onset and correlated version changes), a before/after metric table, trajectory patterns, and tool mix.
  • Compare: pick a version field (e.g. prompt_version) and two values to get a behavioral diff between them. Think of it as git diff for agent behavior.
  • Runs: recent runs with their trajectories. Click one to open a span timeline, where you can inspect each span's inputs, outputs, tokens, cost, and metadata.
  • Versions: every version value seen, with run counts and first/last-seen times.

How drift is detected

Every check is simple enough to reason about. The engine is ~600 lines of dependency-free Python in server/agentdrift_server/engine.

A metric is flagged only when the change is both large and statistically unlikely to be noise:

Check Metrics Test Flagged when
Numeric trajectory length, LLM calls, tool calls, latency, tokens, cost, eval score Welch z-score of the difference in means |z| ≥ 3 and relative change ≥ threshold (e.g. 15% for steps, 20% for latency, 3% for eval score)
Rate success rate, runs with errors, runs with retries, runs with blocked actions Two-proportion z-test |z| ≥ 3 and absolute change ≥ threshold (e.g. 2–3 pp)
Tool mix share of calls per tool Jensen–Shannon divergence JSD ≥ 0.05
Trajectory step-transition bigrams (START → web_search, web_search → llm, …) Jensen–Shannon divergence JSD ≥ 0.05
Model models used by LLM spans Set difference A new model appears in ≥ 10% of runs, or a model disappears

Findings are grouped into categories (trajectory, tool, model, latency, error, cost, quality, guardrail), each graded low, medium, or high. The drift score (0–100) is a weighted summary: behavior 30%, quality 25%, performance 20%, cost 15%, safety 10%. It's a convenience for sorting, not the source of truth. The findings are.

Explanation. The engine runs a single change-point search (CUSUM-style) over the most-drifted signal to estimate the onset. It then compares the dominant value of every version field between the baseline and current windows. Changed fields are reported with timing, e.g. "Drift began 6 min after prompt_version 18 first appeared". If no version changed, it says so and points you to external causes: tool/API behavior, retrieval data, or inputs.

Minimum data: 20 baseline runs and 10 current runs. Below that, the agent shows Collecting baseline.

HTTP API

Interactive docs are at http://localhost:8000/docs.

Method Path Description
POST /api/v1/traces Ingest {"traces": [...]} (the SDK wire format). Idempotent on run_id.
GET /api/v1/agents All agents with status, score, and top findings.
GET /api/v1/agents/{name}/drift?window=50&baseline=200 Full rolling drift report with explanation.
GET /api/v1/agents/{name}/compare?field=prompt_version&a=17&b=18 Behavioral diff between two versions.
GET /api/v1/agents/{name}/runs?limit=50&offset=0 Recent runs.
GET /api/v1/agents/{name}/versions Version values with run counts and first/last seen.
GET /api/v1/runs/{run_id} A run with its spans and fingerprint.
GET /api/v1/health Health check.

You don't need the Python SDK. Any language can post traces in this format:

{
  "traces": [{
    "run_id": "run_82931",
    "agent": "research-agent",
    "started_at": "2026-09-16T09:42:00.000Z",
    "ended_at": "2026-09-16T09:42:02.841Z",
    "status": "success",
    "eval_score": 0.91,
    "versions": {"prompt_version": "18", "model": "claude-sonnet-5", "git_commit": "a91bc2f"},
    "spans": [
      {"span_id": "s1", "type": "tool", "name": "web_search", "status": "success",
       "started_at": "2026-09-16T09:42:00.010Z", "ended_at": "2026-09-16T09:42:00.400Z"},
      {"span_id": "s2", "type": "llm", "name": "summarize", "status": "success", "model": "claude-sonnet-5",
       "input_tokens": 3821, "output_tokens": 911, "cost_usd": 0.031,
       "started_at": "2026-09-16T09:42:00.410Z", "ended_at": "2026-09-16T09:42:02.800Z"}
    ]
  }]
}

Configuration

Server

Setting Default Description
DATABASE_URL / --database-url sqlite:///./agentdrift.db Any SQLAlchemy URL. For PostgreSQL: postgresql+psycopg://user:pass@host/db (install with pip install -e "./server[postgres]").
AGENTDRIFT_HOST / --host 127.0.0.1 Bind address.
AGENTDRIFT_PORT / --port 8000 Port.

SDK (via agentdrift.configure(...) or environment variables)

Option Env var Default Description
endpoint AGENTDRIFT_ENDPOINT http://localhost:8000 Server URL.
enabled AGENTDRIFT_DISABLED=1 disables True Turns recording on or off.
versions AGENTDRIFT_<FIELD> {} Version metadata merged into every run.
capture_io – True Records span/run inputs and outputs. Set to False for sensitive data.
max_io_chars – 2000 Truncation limit for captured inputs and outputs.
policy, approval_handler – None Guardrail configuration.
exporter – HTTP A custom callable that receives each trace dict (useful for tests or other backends).

⚠️ Security note: the MVP server has no authentication. By default it binds to localhost only. Don't expose it publicly, and consider capture_io=False if prompts or tool outputs contain sensitive data.

Running with Docker and PostgreSQL

docker compose up --build

This starts PostgreSQL and the server on http://localhost:8000. Point the SDK there as usual.

Project layout

drift/
├── sdk/                          # `agentdrift` Python SDK (no dependencies)
│   └── agentdrift/
│       ├── tracer.py             # runs, spans, context propagation
│       ├── decorators.py         # @observe, @tool, @llm
│       ├── guardrails.py         # Policy, @guard, GuardrailBlocked
│       ├── client.py             # background batched HTTP exporter
│       └── config.py
├── server/                       # `agentdrift-server` (FastAPI + SQLAlchemy)
│   ├── agentdrift_server/
│   │   ├── engine/               # pure-Python drift engine (no I/O)
│   │   │   ├── fingerprint.py    # trace → behavioral fingerprint
│   │   │   ├── detector.py       # baseline vs current, severity, score
│   │   │   ├── explain.py        # onset estimation + correlated version changes
│   │   │   └── stats.py          # z-tests, JSD, change point
│   │   ├── app.py                # HTTP API + dashboard hosting
│   │   ├── services.py           # ingestion and queries
│   │   ├── db.py                 # agents, runs, spans tables
│   │   └── static/               # dashboard (HTML/CSS/JS, no build step)
│   └── Dockerfile
├── examples/research_agent/      # simulated agent that drifts on a prompt change
├── tests/                        # engine, SDK and API tests
├── docs/idea.md                  # product requirements / vision
└── docker-compose.yml

Development

pip install -e "./server[dev]" -e ./sdk
python -m pytest tests

The engine is pure functions over dicts. To experiment with a new detector, add it to engine/detector.py and cover it in tests/test_engine.py. There's a helper there that generates synthetic runs, and a test that confirms stable behavior is not flagged. Keep that test passing: false positives are what kill trust in a drift detector.

Roadmap

This MVP deliberately keeps the architecture boring. Next up, roughly in order (see docs/idea.md for the full vision):

  • Asynchronous ingestion (Redis queue + worker) and persisted baselines and drift events
  • Alerts: dashboard notifications, then webhooks and Slack
  • LLM-as-judge evaluations
  • OpenTelemetry export and ingestion (GenAI semantic conventions)
  • Deployment/change events (e.g. "prompt v18 deployed at 09:39") alongside first-seen times
  • Richer sequence drift: edit distance, loop detection, trajectory clustering
  • CI mode: replay traces for a PR and fail on behavioral regressions
  • Framework integrations (Claude Agent SDK, LangGraph, OpenAI Agents SDK, …)
  • Authentication, projects, and API keys
  • Kubernetes manifests; Prometheus metrics for Agent Drift itself

Out of scope: agent frameworks, agent builders, model hosting, vector databases, memory frameworks, prompt marketplaces, and generic APM.

Contributing

Issues and pull requests are welcome. Good first contributions:

  • A new example agent (coding agent, browser agent, customer-support agent)
  • An integration that auto-instruments a popular agent framework or LLM client
  • A new drift check, with a test showing it catches real drift and stays quiet on stable data
  • Dashboard improvements

Please run python -m pytest tests before opening a PR, and keep changes focused.

License

Apache License 2.0

About

observability platform for detecting when an AI agent's behavior changes over time, even when traditional monitoring says the system is healthy.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages