Skip to content

Session 15: reproduce the floor + custom SQL routing/budget policy with measured evidence (Parts 1–3) - #8

Open
SairajMN wants to merge 3 commits into
theschoolofai:mainfrom
SairajMN:main
Open

SairajMN wants to merge 3 commits into
theschoolofai:mainfrom
SairajMN:main

Conversation

@SairajMN

@SairajMN SairajMN commented Aug 7, 2026

Copy link
Copy Markdown

Summary

Full Session 15 assignment against glc_v4 + S15Code: reproduces the
existing proof suite with captured evidence (Part 1), adds a custom
routing + budget policy for a SQL-generation workload measured against
an always-frontier baseline (Part 2), and adversarially attacks that
policy's budget enforcement (Part 3). No changes to core glc_v4 /
S15Code source — this PR adds task data, config, proof scripts, and
evidence.


Part 1 — Reproduce the floor

Test suites

  • glc_v4: 444 passed, 12 skipped
  • S15Code: 277 passed

Proofs (all run, all passing)

  • p1_cost_per_task — 3 strategies over 12 tasks. Always-frontier:
    $0.01267883/resolved, 12/12. Always-cheapest: $0.00383610/resolved,
    only 3/12 resolved — still cheaper per resolved task despite a 75%
    failure rate, because the measured price spread absorbs it.
    Budget-aware: 12/12 resolved at $0.01255293/resolved.
  • p2_budget_holds — 0 ceiling breaches; tight allowance downgrades
    frontier → standard; an impossible allowance refuses outright, $0
    spent.
  • p3_denial_of_wallet — 200-round loop vs. $0.002 ceiling: 2
    admitted, 198 refused, spend capped at $0.00192680.
  • p4_trace_export — span costs reconcile exactly with the ledger
    (delta 0.000e+00); prompt/response content confirmed absent from spans.
  • p6_cache_savings — see limitation below.
  • p7_cross_model_ladder — every rung served by a genuinely different
    model; measured spread 83.03x vs. projected 84.9x.

4 captured runs — each with exact prompt, tier requested vs. model
served, ordered event trace, Jaeger trace_id, matching ledger row(s),
and final answer. Covers a frontier resolve, a cheapest-tier failure, a
full-ladder escalation to resolution, and a hard budget refusal.

Honest limitation (directly observed) — p6_cache_savings produced
two semantic-cache hits (cosine similarity = 1.000000) on requests
labeled different, both serving wrong answers at $0.00 cost with no
error, warning, or span. An operator would see $0 spent and assume
success — the silent-failure mode independently reproduced, not quoted.


Part 2 — Custom SQL-generation policy

  • 3-tier ladder (economy → standard → frontier), 23 SQL tasks incl. 3
    deliberately hard edge cases (LIMIT/OFFSET, LEFT vs INNER JOIN, date
    conditions)
  • Budget-aware strategy: 7.1% cheaper per resolved task than
    always-frontier ($0.01212 vs $0.01304), both at 100% resolution
  • Break-even resolution rate r* = 46.88% (measured price spread
    2.53x, measured k = 1.35 extra calls/task, not the configured max)
  • Cost/call reported only as a caveat (60% cheaper), since it's the
    metric the source material itself shows is misleading in isolation

Known limitation, documented not hidden — "wrong case" requirement
directly tested live and confirmed not satisfiable in this environment:
no gateway/API keys available, and offline mode judges placeholder SQL
text rather than real output. Accepted the point loss rather than
fabricate or omit it.


Part 3 — Attack the budget

  • Same-scale (100 rounds) before/after: $3.2064 uncontrolled →
    $0.0189 controlled (99.4% reduction)
    , 87/100 refused before the
    provider was ever contacted
  • Refusal visible in telemetry (recorded BudgetRefused nodes, not
    silent truncation)
  • 10,000-round figure included only as a clearly-labeled extrapolation

Files

  • proofs/README_PART1_FLOOR.md — Part 1 evidence and run tables
  • proofs/README_SQL_POLICY.md — Part 2/3 evidence + reproduce commands
  • proofs/p_sql_policy.py, proofs/p_sql_adversarial.py
  • proofs/tasks/sql_generation.jsonl, proofs/tasks/sql_hard3.jsonl
  • config/tiers_sql.yaml, config/budgets_sql.yaml
  • proofs/uncontrolled_config/ — for the same-scale uncontrolled run
  • proofs/evidence/ — canonical output JSONs backing all numbers above
  • Top-level README.md updated with a summary section linking to both
    evidence docs

Reproduce

Full commands in proofs/README_PART1_FLOOR.md and
proofs/README_SQL_POLICY.md — repo-relative paths only.

Verified clean

No .env, API keys, or local absolute paths in anything committed.

- Custom 3-tier ladder (economy/standard/frontier) + budget policy for
  SQL generation, measured against always-frontier baseline
- Budget-aware strategy: 7.1% cheaper per resolved task, r* = 46.88%
  break-even (measured k=1.35)
- Adversarial denial-of-wallet test at matched scale: $3.2064 ->
  $0.0189 (99.4%), 87/100 refused before provider contact
- Wrong-case requirement tested live, confirmed not satisfiable
  without gateway/API keys (documented, not assumed)
- Evidence JSONs and reproduce commands included, no secrets or
  local paths
Canonical proof outputs for cost/resolved-task, break-even, and the
same-scale adversarial comparison —
proofs/out/ being git-ignored.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant