Meter what a run and a gate actually cost, so the window ceilings hold - #13
Open
ranjitha13g wants to merge 1 commit into
Open
ranjitha13g wants to merge 1 commit into
ranjitha13g wants to merge 1 commit into
Conversation
daily_budget and daily_triage_budget could never fire: both accumulated 0.00 forever, and the morning report's two cost lines were permanently zero. The governor was correct; the engine fed it zero, because it read money from fields only its own test doubles wrote. Three facets, one root cause. engine.py read result['spend_usd'], but AgentRuntime.run() reports budget.spent and never emitted that key. _spend_of read reply['metered_calls'], but GatewayClient.complete -- the callable the HTTP event route passes in -- carries no meter at all. And it looked for cost_usd/usd, while BudgetedGateway writes Charge.as_dict(), whose money field is cost. Fix: runtime.run() publishes spend_usd at the top level; the engine falls back to budget.spent; _spend_of also reads cost, and prices an unmetered reply's own tokens against config/pricing.yaml via the existing Pricing.cost(). No new price appears in Python. The runtime double now builds its return from a real RunBudget.snapshot() instead of a hand-written dict, so this class of drift fails the suite rather than passing it. Reverting the two source files fails 5 tests, two of which already shipped green. 357 passed, up from 353. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Owner
Owner
Session 16 — regraded ✅Score: +100. The verdict below stands: this is a duplicate of an earlier filing (or, for glc_v5 #23, an enhancement rather than a defect), and it is recorded as such. What has changed is the credit. On review, the work here was genuinely done: the bug was found independently, the reproduction is real and the fix is sound. Losing a race you had no way of seeing is not a reason to earn nothing, so this is credited at the full 100 even though it does not carry the first-to-file claim. Original assessment, unchanged: Duplicate of two earlier findings at once: metering the gate is #7 and metering the run is #10, both filed 2026-08-10 against this one's 08-13. Combining them into one PR does not create a new finding. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
daily_budget and daily_triage_budget could never fire: both accumulated 0.00 forever, and the morning report's two cost lines were permanently zero. The governor was correct; the engine fed it zero, because it read money from fields only its own test doubles wrote.
Three facets, one root cause. engine.py read result['spend_usd'], but AgentRuntime.run() reports budget.spent and never emitted that key. _spend_of read reply['metered_calls'], but GatewayClient.complete -- the callable the HTTP event route passes in -- carries no meter at all. And it looked for cost_usd/usd, while BudgetedGateway writes Charge.as_dict(), whose money field is cost.
Fix: runtime.run() publishes spend_usd at the top level; the engine falls back to budget.spent; _spend_of also reads cost, and prices an unmetered reply's own tokens against config/pricing.yaml via the existing Pricing.cost(). No new price appears in Python.
The runtime double now builds its return from a real RunBudget.snapshot() instead of a hand-written dict, so this class of drift fails the suite rather than passing it. Reverting the two source files fails 5 tests, two of which already shipped green.
357 passed, up from 353.