Skip to content

About

A durable, typed runtime for multi-step LLM agents: typed state machine over an append-only event log, resume and deterministic replay, four stop conditions, pre-call budget enforcement, schema-validated tools with permissions, and a regression suite of scripted failure scenarios. 243 tests, no network.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

agent-runtime-reference

CI License: MIT Python 3.11 | 3.12 Tests

The hard part of a multi-step LLM agent is not the prompt. It is state, failure and cost. A demo agent is a while-loop around a chat completion; a real one has to answer harder questions: what happens when the process dies halfway through, what happens when the model returns malformed JSON for the third time, what stops the loop before it spends a hundred dollars, what happens when a tool succeeds but the step after it fails, and how do you prove the behaviour did not change after a refactor. This repository is a small, readable reference implementation of those answers: a typed state machine over an append-only event log, with bounded retries, explicit stop conditions, a budget that is checked before each model call, schema-validated tools with permissions, and a regression harness of scripted failure scenarios. Everything is deterministic and runs offline: no API key, no network, no mocking framework.

What this is, and what it is not

It is a reference implementation, written to be read. Every module is a few hundred lines, every design decision is documented with the alternative that was rejected and what the choice costs. It is a good starting point for a system you are going to write yourself, and a concrete artefact for discussing agent reliability.

It is not a framework. There is no plugin system, no DSL, no configuration language, no scheduler and no server. You wire the pieces together in Python.

It is not a deployment. Nothing here has been run in production, there are no usage numbers behind it, and the price table is sample data that must be re-checked against the provider before anyone bills anything to it.

It is not a model-quality eval. The evals/ suite pins down behaviour under failure (retries, budgets, crashes, replay). Measuring how well a model extracts an invoice needs labelled data and a different harness.

Quickstart

cd agent-runtime-reference
pip install -e ".[dev]"

python examples/document_extraction.py    # a complete 3-step agent, no API key
python -m agent_runtime.evals             # the regression suite
pytest -q                                 # 243 tests
ruff check .

A minimal agent looks like this:

from agent_runtime import Budget, Message, Runtime, ScriptedProvider, StepContext, StepResult

def classify(ctx: StepContext) -> StepResult:
    reply = ctx.complete(
        [Message(role="user", content=str(ctx.data["document"]))],
        prompt_version="classify.v1",
    )
    return ctx.continue_to("finish", {"label": reply.text})

def finish(ctx: StepContext) -> StepResult:
    return ctx.finish({"done": True})

runtime = Runtime(
    {"classify": classify, "finish": finish},
    provider=ScriptedProvider(["invoice"]),          # swap for OpenAICompatibleProvider
    budget=Budget(max_model_calls=10, max_cost_usd=0.50),
)

state = runtime.start_and_run("run-1", {"document": "INVOICE 42"}, "classify")
print(state.status.value, state.data, state.usage.cost_usd)

runtime.resume("run-1")        # rebuild the state from the event log
runtime.replay("run-1")        # re-execute against recorded inputs and compare

Pointing it at a real endpoint is a one-line change:

from agent_runtime import OpenAICompatibleProvider

provider = OpenAICompatibleProvider("https://api.example.com/v1", api_key=os.environ["API_KEY"])

Design decisions and their trade-offs

Event sourcing, not a state row. RunState is never written to storage. What is written is an ordered list of events, and the state is recomputed by folding them. Resume is then just "read the log and fold it", with no half-written row to reconcile after a crash, and replay is possible at all because the recorded model responses and tool results are still there. The cost: reading a run is O(events), the log grows without compaction, and changing an event's shape requires a migration or a versioned reader. For agent runs, which are short and few, that trade is easy. For a high-volume system it would need snapshots.

Idempotency keys derived from the input, not the output. Each step gets a deterministic key from the run id, the step position, the step name and the data it is about to read. A step whose key is already in the state is a no-op, both in the runtime and in the fold, which makes an at-least-once writer safe: a retried append cannot double count tokens or advance the run twice. The cost: the key covers the whole input dict, so two logically identical inputs that differ in a single irrelevant field produce different keys, and a step that reads something outside the state is not covered at all.

The budget is checked before the call, not after. The ledger builds a deliberately pessimistic estimate of the next call, prompt tokens plus the full output allowance, and refuses the call if that estimate would cross a limit. Checking afterwards only tells you the money is already gone. The cost: the prompt estimate is a characters-per-token heuristic rather than a real tokenizer, so a run can stop slightly earlier than strictly necessary. That is the right direction to be wrong in for a hard ceiling, and the authoritative token counts always come back from the provider and are what the ledger records.

Self-reported confidence is rejected as a routing signal. A model's "confidence": 0.95 is a token sequence, not a probability: it is uncalibrated, it is correlated with the very failure it is supposed to detect (a model that misreads a field is usually just as sure about the misreading), and it moves silently whenever the prompt, the temperature or the model version changes. guardrails.py routes on evidence instead: did it validate against the schema, does the value appear verbatim in the source, does it exist in reference data, did any cross-field rule fail, did the model need a repair round, do independent samples agree. The cost: those signals have to be computed, they are domain-specific, and the weights combining them are a starting point that should be refitted against real review outcomes rather than trusted as given.

Failure is a value, not an exception. Tool validation errors, permission refusals and tool exceptions all come back as a ToolResult carrying a ToolError, so the agent can see them and correct itself. Exhausted retries, an exhausted budget and a permanent tool failure end the run as DEGRADED with a machine-readable reason and whatever partial output was produced. Only an unexpected exception inside a step, which is a bug, ends as FAILED. The cost: step authors have to check results instead of relying on exceptions to propagate, which is more code in each step.

One retry for structured output, not a loop. If the model returns something that does not fit the schema, the validation error is fed back once. If the second attempt still fails, the run degrades. The cost: a model that would have succeeded on the third try does not get one. That is deliberate, because at that point the problem is the prompt or the schema and more attempts only make the incident more expensive.

A step that has run a non-idempotent tool is never retried. Retrying a step that already booked a payment would book it twice, so the runtime degrades the run instead and lets a human look at it. The cost: some genuinely transient failures turn into degraded runs that a smarter compensation mechanism could have recovered.

Modules

Module What it holds
contracts.py The typed vocabulary: RunState (versioned), StepResult, ToolCall, ToolResult, Budget, Usage, TraceEvent, and the transition union ContinueTo / Finish / RetryStep. Frozen models, discriminated unions, and the two derivation functions (idempotency key, progress marker).
runtime.py The loop: typed state machine, bounded exponential backoff, four stop conditions, degraded terminal states, idempotency, and replay. StepContext is the only door a step has to the model, the tools and the budget.
persistence.py Store protocol, InMemoryStore and SQLiteStore, the pure apply_event fold, resume and replay. Append-only, with the sequence number doubling as an optimistic-concurrency token.
tools.py Tool registry: JSON Schema derived from a Pydantic model, permission levels (READ, WRITE, DANGEROUS), an idempotency flag, argument validation before execution, and an approval callback for dangerous tools.
budget.py Price table as data, a token estimate for pre-call checks, and a ledger with limits on tokens, cost and call count that can be seeded from a resumed run.
providers.py ModelProvider protocol, ScriptedProvider (used by every test), ReplayProvider, OpenAICompatibleProvider (injectable transport), and structured output with a single repair round.
guardrails.py Cross-field rules, a reference-data hook, and a three-tier routing policy (auto-accept, escalate with a prefilled value, leave to a human) driven by evidence rather than self-reported confidence.
tracing.py A structured trace of every step and model call (prompt version, model, tokens, cost, latency, retries, outcome) with export_jsonl. Kept separate from the event log on purpose.
evals/ The regression harness: agent.py (the agent under test), scenarios.py (16 scenarios, 92 invariants), runner.py (pass/fail per invariant, CI exit code).

Running the evals

python -m agent_runtime.evals              # table of scenarios
python -m agent_runtime.evals -v           # list passing invariants too
python -m agent_runtime.evals --json       # machine-readable report
python -m agent_runtime.evals -k budget    # only scenarios whose name matches

The suite covers the happy path, tool argument validation failure, invalid JSON followed by a successful repair, an unrepairable reply, budget exhaustion by call count and by tokens, a retried provider error, a retried tool error, a permanent tool failure, a simulated crash followed by resume in a new process, replay determinism, a dangerous tool with and without approval, the no-progress detector, the step limit, and a transient failure after a non-idempotent write. Each scenario asserts named invariants, so a failure report says which property broke rather than only that something did.

The same scenarios also run under pytest (tests/test_evals.py), one pytest case per scenario, so a broken invariant fails the normal test run too.

Known limitations

  • No concurrency. One runtime executes one run at a time, and SQLiteStore uses a single connection behind a lock. Parallelism belongs one level up (one runtime per worker); the sequence-number collision is the only protection against two writers on the same run.
  • No snapshots or compaction. Resuming a long run means folding every event. Fine for runs of tens of steps, not for thousands.
  • RunState.data is a plain dict. The model is frozen, but the dict inside it is not deep-immutable, so a step could mutate it in place. The convention is to return changes as StepResult.output; that is a convention, not a guarantee.
  • Token estimates are a heuristic. estimate_tokens divides characters by four. It is only used for the pre-call budget check, and it is deliberately not a vendor tokenizer.
  • The price table is sample data. Prices change and differ by region and contract. Load your own with PriceTable.from_json_file.
  • Guardrail weights are not fitted. SIGNAL_WEIGHTS and the policy thresholds are a defensible starting point, not a result. Refit them on labelled review outcomes.
  • One repair attempt is hardcoded as the policy. It is a deliberate choice, not a configurable one, and a different domain may want a different rule.
  • Only one real provider adapter. OpenAICompatibleProvider covers endpoints that speak the chat-completions shape. Anything else needs a new adapter against the same protocol.

License

MIT. See LICENSE.

About

A durable, typed runtime for multi-step LLM agents: typed state machine over an append-only event log, resume and deterministic replay, four stop conditions, pre-call budget enforcement, schema-validated tools with permissions, and a regression suite of scripted failure scenarios. 243 tests, no network.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages