Conversation
- Custom 3-tier ladder (economy/standard/frontier) + budget policy for SQL generation, measured against always-frontier baseline - Budget-aware strategy: 7.1% cheaper per resolved task, r* = 46.88% break-even (measured k=1.35) - Adversarial denial-of-wallet test at matched scale: $3.2064 -> $0.0189 (99.4%), 87/100 refused before provider contact - Wrong-case requirement tested live, confirmed not satisfiable without gateway/API keys (documented, not assumed) - Evidence JSONs and reproduce commands included, no secrets or local paths
Canonical proof outputs for cost/resolved-task, break-even, and the same-scale adversarial comparison — proofs/out/ being git-ignored.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Full Session 15 assignment against
glc_v4+S15Code: reproduces theexisting proof suite with captured evidence (Part 1), adds a custom
routing + budget policy for a SQL-generation workload measured against
an always-frontier baseline (Part 2), and adversarially attacks that
policy's budget enforcement (Part 3). No changes to core
glc_v4/S15Codesource — this PR adds task data, config, proof scripts, andevidence.
Part 1 — Reproduce the floor
Test suites
glc_v4: 444 passed, 12 skippedS15Code: 277 passedProofs (all run, all passing)
p1_cost_per_task— 3 strategies over 12 tasks. Always-frontier:$0.01267883/resolved, 12/12. Always-cheapest: $0.00383610/resolved,
only 3/12 resolved — still cheaper per resolved task despite a 75%
failure rate, because the measured price spread absorbs it.
Budget-aware: 12/12 resolved at $0.01255293/resolved.
p2_budget_holds— 0 ceiling breaches; tight allowance downgradesfrontier → standard; an impossible allowance refuses outright, $0
spent.
p3_denial_of_wallet— 200-round loop vs. $0.002 ceiling: 2admitted, 198 refused, spend capped at $0.00192680.
p4_trace_export— span costs reconcile exactly with the ledger(delta 0.000e+00); prompt/response content confirmed absent from spans.
p6_cache_savings— see limitation below.p7_cross_model_ladder— every rung served by a genuinely differentmodel; measured spread 83.03x vs. projected 84.9x.
4 captured runs — each with exact prompt, tier requested vs. model
served, ordered event trace, Jaeger
trace_id, matching ledger row(s),and final answer. Covers a frontier resolve, a cheapest-tier failure, a
full-ladder escalation to resolution, and a hard budget refusal.
Honest limitation (directly observed) —
p6_cache_savingsproducedtwo semantic-cache hits (cosine similarity = 1.000000) on requests
labeled different, both serving wrong answers at $0.00 cost with no
error, warning, or span. An operator would see $0 spent and assume
success — the silent-failure mode independently reproduced, not quoted.
Part 2 — Custom SQL-generation policy
deliberately hard edge cases (LIMIT/OFFSET, LEFT vs INNER JOIN, date
conditions)
always-frontier ($0.01212 vs $0.01304), both at 100% resolution
2.53x, measured k = 1.35 extra calls/task, not the configured max)
metric the source material itself shows is misleading in isolation
Known limitation, documented not hidden — "wrong case" requirement
directly tested live and confirmed not satisfiable in this environment:
no gateway/API keys available, and offline mode judges placeholder SQL
text rather than real output. Accepted the point loss rather than
fabricate or omit it.
Part 3 — Attack the budget
$0.0189 controlled (99.4% reduction), 87/100 refused before the
provider was ever contacted
BudgetRefusednodes, notsilent truncation)
Files
proofs/README_PART1_FLOOR.md— Part 1 evidence and run tablesproofs/README_SQL_POLICY.md— Part 2/3 evidence + reproduce commandsproofs/p_sql_policy.py,proofs/p_sql_adversarial.pyproofs/tasks/sql_generation.jsonl,proofs/tasks/sql_hard3.jsonlconfig/tiers_sql.yaml,config/budgets_sql.yamlproofs/uncontrolled_config/— for the same-scale uncontrolled runproofs/evidence/— canonical output JSONs backing all numbers aboveREADME.mdupdated with a summary section linking to bothevidence docs
Reproduce
Full commands in
proofs/README_PART1_FLOOR.mdandproofs/README_SQL_POLICY.md— repo-relative paths only.Verified clean
No
.env, API keys, or local absolute paths in anything committed.