Skip to content

Record what a run actually spent, so daily_budget can bind - #10

Open
Sujthr wants to merge 1 commit into
theschoolofai:mainfrom
Sujthr:fix/run-spend-never-recorded
Open

Sujthr wants to merge 1 commit into
theschoolofai:mainfrom
Sujthr:fix/run-spend-never-recorded

Conversation

@Sujthr

@Sujthr Sujthr commented Aug 10, 2026

Copy link
Copy Markdown

AutonomousEventEngine records a completed run's cost with:

usd=float(result.get("spend_usd") or 0.0)

Runtime.run() never returns spend_usd. Its result dict is {"run_id", "status", "answer", "provider", "model", "graph", "trace", "events", "principal", "budget", "economics", "allocations"}, and the figure lives at
result["budget"]["spent"] — run_budget.snapshot().

So every run is recorded at $0.00 and window_spend(..., kind="run") stays zero
for the life of the installation.

Three consequences, one root cause

  1. daily_budget can never bind. admit_run compares the window's run spend
    against it, and that spend is always 0. Section 10 makes this the ceiling that
    survives "escape the budget by starting more runs": a per-run ceiling bounds
    nothing for an agent that can start more of them.

  2. A run's ceiling is never squeezed by the window. admit_run computes
    remaining = daily_budget - spent and hands the run min(per_run, remaining).
    With spent pinned at 0, remaining is always the full daily budget, so the
    last run of the day is handed as much as the first. Section 10 asks for "each
    run's ceiling shrunk by whatever the window has left".

  3. The morning report understates the bill. Observed live on 2026-08-10 after
    four real runs against Gemini:

    - acted: 4
    - cost of watching: $0.00066000
    - cost of doing:    $0.00000000
    

    Section 9 makes that report the thing an operator reads after an unattended
    night, and it showed a night that cost nothing.

max_runs_per_day still bounds the count, so this is not unbounded spend — but
the money ceiling specifically is inert.

Why the suite did not catch it

tests/test_autonomy_governor.py defines:

class _Runtime:
    async def run(self, *, prompt, **_):
        return {"run_id": ..., "status": "completed", "spend_usd": self.spend}

The stub invents the key production never sends, so
test_a_daily_budget_caps_the_window_and_shrinks_the_last_run passes while the
control it covers cannot fire — Section 12's "a green check on a control that
never ran is worse than a red one".

tests/test_run_spend_is_recorded.py uses a stub that mirrors the real
Runtime.run contract instead, including the absence of spend_usd. Its 3 tests
fail before this change and pass after; the existing 10 governor tests still pass,
because spend_usd is still honoured when a caller supplies it.

`AutonomousEventEngine` records a completed run's cost with:

    usd=float(result.get("spend_usd") or 0.0)

`Runtime.run()` never returns `spend_usd`. Its result dict is `{"run_id",
"status", "answer", "provider", "model", "graph", "trace", "events", "principal",
"budget", "economics", "allocations"}`, and the figure lives at
`result["budget"]["spent"]` — `run_budget.snapshot()`.

So every run is recorded at $0.00 and `window_spend(..., kind="run")` stays zero
for the life of the installation.

## Three consequences, one root cause

1. **`daily_budget` can never bind.** `admit_run` compares the window's run spend
   against it, and that spend is always 0. Section 10 makes this the ceiling that
   survives "escape the budget by starting more runs": a per-run ceiling bounds
   nothing for an agent that can start more of them.

2. **A run's ceiling is never squeezed by the window.** `admit_run` computes
   `remaining = daily_budget - spent` and hands the run `min(per_run, remaining)`.
   With `spent` pinned at 0, `remaining` is always the full daily budget, so the
   last run of the day is handed as much as the first. Section 10 asks for "each
   run's ceiling shrunk by whatever the window has left".

3. **The morning report understates the bill.** Observed live on 2026-08-10 after
   four real runs against Gemini:

       - acted: 4
       - cost of watching: $0.00066000
       - cost of doing:    $0.00000000

   Section 9 makes that report the thing an operator reads after an unattended
   night, and it showed a night that cost nothing.

`max_runs_per_day` still bounds the count, so this is not unbounded spend — but
the money ceiling specifically is inert.

## Why the suite did not catch it

`tests/test_autonomy_governor.py` defines:

    class _Runtime:
        async def run(self, *, prompt, **_):
            return {"run_id": ..., "status": "completed", "spend_usd": self.spend}

The stub invents the key production never sends, so
`test_a_daily_budget_caps_the_window_and_shrinks_the_last_run` passes while the
control it covers cannot fire — Section 12's "a green check on a control that
never ran is worse than a red one".

`tests/test_run_spend_is_recorded.py` uses a stub that mirrors the real
`Runtime.run` contract instead, including the absence of `spend_usd`. Its 3 tests
fail before this change and pass after; the existing 10 governor tests still pass,
because `spend_usd` is still honoured when a caller supplies it.
@theschoolofai

Copy link
Copy Markdown
Owner

Session 16 — graded ✅

Score: +100.

A run's actual spend was never recorded, so daily_budget had nothing to bind against. Distinct from #7: that meters the gate, this meters the run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants