Conversation
…ttempt-guard Measured live against Groq + Gemini through glc_v4, with Jaeger traces fetched back by ID. Full evidence in the README's "Evidence" section. Part 1 - reproduce the floor Both suites green (glc_v4 445 passed; S15Code 298). All five proofs pass live. proofs/repro_floor_runs.py captures prompt, tier/model, ordered event trace, Jaeger trace ID, ledger rows and final answer for four runs - only p4 among the shipped proofs exports telemetry itself, so none of them hands back a trace ID. The limitation the traces exposed: a well-formed trace cannot tell you a call never happened. A rate-limited run emitted a structurally perfect 9-span tree whose costs summed correctly to $0.00. Cost telemetry needs a denominator - calls attempted, not just calls made. Part 2 - my task class, ladder and budget policy proofs/tasks/support_tickets.jsonl: 18 customer-support tickets (8 trivial, 6 moderate, 4 hard) asking for a ready-to-send reply draft. config/support_desk/: a self-contained ladder - groq/gpt-oss-20b < groq/gpt-oss-120b < gemini/gemini-3.1-flash-lite - plus budget policy and a judge panel disjoint from every rung. Kept out of the shipped config/ so the shipped ladder is untouched. Measured: the budget-aware cascade costs $0.00018922 per resolved task at 18/18 resolved, against the always-frontier baseline's $0.00027514 at 18/18 - 31.2% cheaper for the same resolution rate. Break-even for the cheap rung is 85.9% against the frontier baseline and 99.2% against the cascade. Every model name was verified callable against the live APIs first: the shipped config names gemini-3.1-pro (404) and gemini-3.1-flash (not served), and free-tier Gemini grants no quota at all on pro-class models. Recorded in the README rather than silently substituted. Part 3 - attack the budget s15code/economics/attempt_guard.py closes a real gap: MeteredTransport does not charge a call that raises, on the stated assumption that it "consumed no tokens the gateway could report". A provider that generates output and only then drops the connection breaks that assumption, and because RunBudget.calls only increments inside charge(), the call ceiling never trips either - so an adversarial retry loop runs unbounded with the ledger reporting $0.00 while a real provider invoices $1.09. With the guard, 90 of 150 asks are refused pre-flight and the true bill is bounded by the attempt ceiling. Refusals are visible in exported telemetry (150 ERROR-status spans, 90 BudgetRefused nodes), both traces fetched back from Jaeger. Also fixed, because it silently corrupted the headline number: s15code/economics/patient_transport.py. evals.yaml paces and retries the JUDGE path on the grounds that a rate limit is a transport failure and "giving up and calling the task unresolved would not be" honest - but that reasoning was never applied to the ANSWER path, where a 429 became an empty answer the judge then scored unresolved. A provider's per-minute quota was entering the resolution rate, the denominator of cost-per-resolved-task. Retries transient failures with backoff, raises permanent ones immediately, never fabricates a response. Live mode only; offline CI unchanged. The one case the policy got wrong: on s12_two_issues_one_ticket the cascade opened cheap, failed, escalated, and paid $0.00060875 where the frontier baseline paid $0.00032675 - 86% more on that task. tests/test_attempt_guard.py and tests/test_patient_transport.py cover both new modules hermetically. No edit to the shipped controller.py or budget.py. No secrets, no .env.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Evidence section is at the end of README.md: the workload and ladder, cost per call vs cost per resolved task against an always-frontier baseline, the break-even resolution rate, Jaeger traces with per-span costs, the adversarial test failing before the control and refused after, and one case the policy got wrong.