Session 15 of EAGV3. Hand a run a ceiling and a principal, and the agent plans against them: every node declares the capability tier it needs, the allowance is re-divided across the live frontier on every planning round, and a controller admits, downgrades, branches or refuses each call before it is made. The same durable event journal the graph already writes is then exported as OpenTelemetry spans, with token usage and cost per span.
Two things are true no matter what the model does with its tokens:
- No call is made without being metered. The controller owns the transport; there is no code path to a provider that skips the ledger.
- No run spends past its ceiling. Enforcement is deterministic code. A budget asked for in a prompt leaks — models are token-elastic (TALE, 2026), so they sail past a tight ceiling while sincerely agreeing to it.
S14Code shipped two packages side by side, with the session's own work hidden
inside the previous session's namespace. This repo has exactly one importable
package, and nothing is nested inside a prior session's name.
S15Code/
├── s15code/
│ ├── main.py FastAPI app factory
│ ├── cli.py `s15code serve`
│ ├── routes.py runs, facts, documents, memory search, trace
│ ├── a2a_routes.py agent card + JSON-RPC
│ ├── gateway.py the gateway client (never holds a credential)
│ ├── planner.py the constrained GraphPatch proposal boundary
│ ├── runtime.py one request through the live graph
│ ├── tools.py the small non-browser skill surface
│ ├── core/
│ │ ├── live_graph/ executor, durable event journal, patches (from S13)
│ │ ├── memory/ typed scoped memory, semantic chunking (from S13)
│ │ └── a2a/ the agent-to-agent boundary (from S13)
│ ├── ui/ catalog, validator, surface, AG-UI, HITL (from S14)
│ ├── economics/ NEW — budget-aware planning
│ │ ├── config.py loads the three YAML files
│ │ ├── pricing.py per-model prices
│ │ ├── tiers.py the capability ladder; a node declares a tier
│ │ ├── budget.py allowance, spend, reservations, allocation
│ │ ├── policy.py proceed / downgrade / branch / refuse
│ │ └── controller.py the hard controller at the call seam
│ ├── telemetry/ NEW — the same journal, as OTel spans
│ │ └── spans.py run → agent loop → plan → node → provider call
│ └── evals/ NEW — did the answer RESOLVE the task?
│ ├── config.py the rubric, the bar and the judge panel, from YAML
│ ├── judge.py LLM-as-judge on a generic, task-agnostic rubric
│ └── tasks.py reads a task set; never contains one
├── config/ tiers.yaml · pricing.yaml · budgets.yaml · evals.yaml
├── proofs/ the generic proof harness
│ └── tasks/ task sets, as DATA a reviewer can replace
└── tests/
The graph writes one durable journal. It now has three consumers, and no parallel event system exists:
| Consumer | Reads the journal as |
|---|---|
| the executor | graph replay and crash recovery (S13) |
s15code.ui.agui |
AG-UI events for a browser (S14) |
s15code.telemetry.spans |
OpenTelemetry spans for a collector (S15) |
The controller writes each metered call into the node's own result, so the
journal carries tokens, cost, tier and the budget decision. Delete the
materialised graph and the trace still builds from the tape alone — p4 checks
exactly that.
No tier name, model, provider, price, threshold or budget appears in Python.
| Decision | Lives in |
|---|---|
| what tiers exist, and what each expands to on the wire | config/tiers.yaml |
| which tier a graph role asks for | config/tiers.yaml → role_tiers |
| what a model costs, and cache-read discounts | config/pricing.yaml |
| default allowance, per-principal caps | config/budgets.yaml |
| downgrade / refuse ratios, reserve, call ceilings, token estimation | config/budgets.yaml |
config/ is resolved from S15_CONFIG_DIR when set, otherwise from beside the
package. The unit tests build their own ladder with invented tier names, which
is the real check that the library never depends on the shipped ones.
uv sync
uv run pytest -q
uv run ruff check .
uv run s15code serve # http://127.0.0.1:8113The gateway is a separate process on http://127.0.0.1:8111 (GLC_BASE_URL).
S15Code holds no provider credential; copy .env.example to .env and set paths,
never keys.
A budgeted run over HTTP:
curl -s localhost:8113/v1/agent/runs -H 'content-type: application/json' -d '{
"tenant_id": "acme", "project_id": "research", "user_id": "rohan",
"prompt": "<any task>", "budget": 0.02
}' | jq '.budget'The response carries a budget ledger (total, spent, remaining, pressure, every
charge, every refusal), the allocations the planner made each round, and the
tier each node declared. GET /v1/agent/runs/{id}/trace returns the same run as
a span tree.
Omit budget and the run behaves exactly as it did before economics existed —
the layer is additive.
One harness, one code path, six proofs. Each takes the task (or task set, or
pair set), budget and principal as arguments, asserts real invariants, exits
non-zero on failure, and writes JSON to proofs/out/.
uv run python proofs/p1_cost_per_task.py --tasks proofs/tasks/mixed.jsonl
uv run python proofs/p2_budget_holds.py --task "<any task>" --budget 0.02
uv run python proofs/p3_denial_of_wallet.py --task "<any task>" --budget 0.002
uv run python proofs/p4_trace_export.py --task "<any task>" --budget 0.02
uv run python proofs/p6_cache_savings.py --pairs proofs/pairs/paraphrases.jsonl
uv run python proofs/p7_cross_model_ladder.py --task "<any task>"| Proof | Proves |
|---|---|
p1_cost_per_task |
The same task set through always-frontier, always-cheapest-with-retries and the budget-aware cascade, on one ledger. Reports spend, calls, cost per call and cost per resolved task per strategy, with "resolved" decided by a generic rubric judge (s15code.evals) and never by a per-task answer key. Whether the cheapest rung shows the signature failure mode — lower cost per call, higher cost per resolved task — is reported as a finding, not asserted. |
p2_budget_holds |
A run given a ceiling stays under it at every ceiling; a tight allowance downgrades the tier a node asked for; an unaffordable ceiling refuses instead of overspending; provider calls and ledger entries agree exactly. |
p3_denial_of_wallet |
An adversarial planner that earns one more node from every outcome, forever, cannot spend past the ceiling. Refusals are visible graph failures. Reports the bill the same loop would have run up uncontrolled. |
p4_trace_export |
The journal exports as run → agent loop → plan → node → provider call, with gen_ai.usage.* and cost on every provider-call span, summing exactly to the ledger. Content capture off. Works with no collector. |
p6_cache_savings |
The gateway's semantic cache, against the real embedder. A 768-dim nomic vector is confirmed to be neither a stub nor a constant; the similarity the gateway acts on is checked against a cosine computed independently; a hit is billed $0 and its saving is read off the cold call it replaced. Then the threshold is swept over a labelled pair set (proofs/pairs/), reporting true- and false-positive rates per threshold and per negative family. Whether a collision-free threshold exists is a finding, not an assertion — and on the shipped set it does not. |
p7_cross_model_ladder |
Every rung is a different model on a different provider; budget pressure walks the whole ladder down, one model at a time; projected cost is monotone; the measured top-to-bottom spread is reported as a multiple. |
Two modes, one code path. If the gateway at --base-url answers, the proofs
make real calls and meter real money. Otherwise (or with --offline) a
deterministic transport stands in — the policy, ladder, budget, journal and span
export are all the real implementation, only the network is replaced. That is what
lets CI run the same proof with no key and no collector.
The ceilings p2 uses to force a downgrade and a refusal are derived from the
configured ladder, not written down, so editing config/tiers.yaml changes the
numbers rather than breaking the proof.
Useful flags: --offline, --respond-as ui, --principal tenant/project/user,
--otel-endpoint http://127.0.0.1:4318/v1/traces, --config-dir, --label,
--live-embeddings.
s15code.telemetry.export_run turns a journal into spans through the real OTel
SDK. Attributes follow the GenAI semantic conventions —
gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model,
gen_ai.usage.input_tokens, gen_ai.usage.output_tokens. Those conventions are
pre-stable and moved to their own repository in June 2026 with no tagged release,
so cost has no blessed attribute yet: it is emitted as s15.cost with
s15.currency, clearly marked as a vendor extension rather than pretending to be
standard.
Set S15_OTEL_EXPORTER_ENDPOINT (or --otel-endpoint) to send the spans to
Jaeger or any OTLP receiver; Jaeger ingests OTLP natively. Leave it unset and the
span tree is still built and assertable in memory while nothing goes over the
wire, so tests need no collector.
Content capture is off by default. Prompts and completions are PII. They are
attached only when a caller passes capture_content=True or sets
S15_OTEL_CAPTURE_CONTENT=1.
The ledger covers gateway model calls — the paid ones. Semantic-memory embeddings run locally against Ollama and never touch the gateway, so they cost nothing per token and are not in the ledger.
Admission prices the worst case: output is bounded by the tier's max_tokens,
which the provider honours, and input by an over-estimate of the prompt actually
being sent (chars_per_token and input_estimate_safety in budgets.yaml). That
makes the projection a real upper bound rather than an average. If a provider
ignored max_tokens the call already in flight could overshoot — so the ledger is
also an absolute stop, and max_calls_per_run / max_calls_per_node bound the
loop even when every price estimate is wrong. A test drives exactly that case.
Carried forward (imports renamed to s15code.*, behaviour unchanged): the
live graph and its journal, scoped memory and semantic chunking, the A2A boundary
and gRPC binding, the A2UI catalog/validator/surface/AG-UI/HITL layer, the
gateway boundary, the deterministic and LLM planners, the non-browser skills.
New in this session: s15code/economics/ (six modules), s15code/telemetry/,
config/ (three files), the budget/principal arguments on a run, the
/trace route, and a rewritten proofs/ harness.
Deliberately dropped: S14's showcase.py and its /dashboard route hardcoded
one use case (a five-paper research corpus, with its title in the code). The
/v1/harness/surface route depended on one specific S14 proof artefact. Both
would have violated the no-hardcoding rule this session is built around.
The generated protobuf modules under s15code/core/a2a/ keep their original
filenames. They are reproduced verbatim because the serialized descriptor is
keyed on the .proto file name, and hand-editing generated gencode is worse than
a stale name.