Conversation
`AutonomousEventEngine` records a completed run's cost with:
usd=float(result.get("spend_usd") or 0.0)
`Runtime.run()` never returns `spend_usd`. Its result dict is `{"run_id",
"status", "answer", "provider", "model", "graph", "trace", "events", "principal",
"budget", "economics", "allocations"}`, and the figure lives at
`result["budget"]["spent"]` — `run_budget.snapshot()`.
So every run is recorded at $0.00 and `window_spend(..., kind="run")` stays zero
for the life of the installation.
## Three consequences, one root cause
1. **`daily_budget` can never bind.** `admit_run` compares the window's run spend
against it, and that spend is always 0. Section 10 makes this the ceiling that
survives "escape the budget by starting more runs": a per-run ceiling bounds
nothing for an agent that can start more of them.
2. **A run's ceiling is never squeezed by the window.** `admit_run` computes
`remaining = daily_budget - spent` and hands the run `min(per_run, remaining)`.
With `spent` pinned at 0, `remaining` is always the full daily budget, so the
last run of the day is handed as much as the first. Section 10 asks for "each
run's ceiling shrunk by whatever the window has left".
3. **The morning report understates the bill.** Observed live on 2026-08-10 after
four real runs against Gemini:
- acted: 4
- cost of watching: $0.00066000
- cost of doing: $0.00000000
Section 9 makes that report the thing an operator reads after an unattended
night, and it showed a night that cost nothing.
`max_runs_per_day` still bounds the count, so this is not unbounded spend — but
the money ceiling specifically is inert.
## Why the suite did not catch it
`tests/test_autonomy_governor.py` defines:
class _Runtime:
async def run(self, *, prompt, **_):
return {"run_id": ..., "status": "completed", "spend_usd": self.spend}
The stub invents the key production never sends, so
`test_a_daily_budget_caps_the_window_and_shrinks_the_last_run` passes while the
control it covers cannot fire — Section 12's "a green check on a control that
never ran is worse than a red one".
`tests/test_run_spend_is_recorded.py` uses a stub that mirrors the real
`Runtime.run` contract instead, including the absence of `spend_usd`. Its 3 tests
fail before this change and pass after; the existing 10 governor tests still pass,
because `spend_usd` is still honoured when a caller supplies it.
Owner
Session 16 — graded ✅Score: +100. A run's actual spend was never recorded, so |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
AutonomousEventEnginerecords a completed run's cost with:Runtime.run()never returnsspend_usd. Its result dict is{"run_id", "status", "answer", "provider", "model", "graph", "trace", "events", "principal", "budget", "economics", "allocations"}, and the figure lives atresult["budget"]["spent"]—run_budget.snapshot().So every run is recorded at $0.00 and
window_spend(..., kind="run")stays zerofor the life of the installation.
Three consequences, one root cause
daily_budgetcan never bind.admit_runcompares the window's run spendagainst it, and that spend is always 0. Section 10 makes this the ceiling that
survives "escape the budget by starting more runs": a per-run ceiling bounds
nothing for an agent that can start more of them.
A run's ceiling is never squeezed by the window.
admit_runcomputesremaining = daily_budget - spentand hands the runmin(per_run, remaining).With
spentpinned at 0,remainingis always the full daily budget, so thelast run of the day is handed as much as the first. Section 10 asks for "each
run's ceiling shrunk by whatever the window has left".
The morning report understates the bill. Observed live on 2026-08-10 after
four real runs against Gemini:
Section 9 makes that report the thing an operator reads after an unattended
night, and it showed a night that cost nothing.
max_runs_per_daystill bounds the count, so this is not unbounded spend — butthe money ceiling specifically is inert.
Why the suite did not catch it
tests/test_autonomy_governor.pydefines:The stub invents the key production never sends, so
test_a_daily_budget_caps_the_window_and_shrinks_the_last_runpasses while thecontrol it covers cannot fire — Section 12's "a green check on a control that
never ran is worse than a red one".
tests/test_run_spend_is_recorded.pyuses a stub that mirrors the realRuntime.runcontract instead, including the absence ofspend_usd. Its 3 testsfail before this change and pass after; the existing 10 governor tests still pass,
because
spend_usdis still honoured when a caller supplies it.