Skip to content

S15: support-desk routing ladder, budget policy, and an adversarial attempt-guard - #19

Open
AaryanG21 wants to merge 1 commit into
theschoolofai:mainfrom
AaryanG21:s15-budget-policy-aaryan
Open

AaryanG21 wants to merge 1 commit into
theschoolofai:mainfrom
AaryanG21:s15-budget-policy-aaryan

Conversation

@AaryanG21

Copy link
Copy Markdown

Evidence section is at the end of README.md: the workload and ladder, cost per call vs cost per resolved task against an always-frontier baseline, the break-even resolution rate, Jaeger traces with per-span costs, the adversarial test failing before the control and refused after, and one case the policy got wrong.

…ttempt-guard

Measured live against Groq + Gemini through glc_v4, with Jaeger traces fetched
back by ID. Full evidence in the README's "Evidence" section.

Part 1 - reproduce the floor
  Both suites green (glc_v4 445 passed; S15Code 298). All five proofs pass
  live. proofs/repro_floor_runs.py captures prompt, tier/model, ordered event
  trace, Jaeger trace ID, ledger rows and final answer for four runs - only p4
  among the shipped proofs exports telemetry itself, so none of them hands back
  a trace ID.
  The limitation the traces exposed: a well-formed trace cannot tell you a call
  never happened. A rate-limited run emitted a structurally perfect 9-span tree
  whose costs summed correctly to $0.00. Cost telemetry needs a denominator -
  calls attempted, not just calls made.

Part 2 - my task class, ladder and budget policy
  proofs/tasks/support_tickets.jsonl: 18 customer-support tickets (8 trivial,
  6 moderate, 4 hard) asking for a ready-to-send reply draft.
  config/support_desk/: a self-contained ladder - groq/gpt-oss-20b <
  groq/gpt-oss-120b < gemini/gemini-3.1-flash-lite - plus budget policy and a
  judge panel disjoint from every rung. Kept out of the shipped config/ so the
  shipped ladder is untouched.
  Measured: the budget-aware cascade costs $0.00018922 per resolved task at
  18/18 resolved, against the always-frontier baseline's $0.00027514 at 18/18 -
  31.2% cheaper for the same resolution rate. Break-even for the cheap rung is
  85.9% against the frontier baseline and 99.2% against the cascade.
  Every model name was verified callable against the live APIs first: the
  shipped config names gemini-3.1-pro (404) and gemini-3.1-flash (not served),
  and free-tier Gemini grants no quota at all on pro-class models. Recorded in
  the README rather than silently substituted.

Part 3 - attack the budget
  s15code/economics/attempt_guard.py closes a real gap: MeteredTransport does
  not charge a call that raises, on the stated assumption that it "consumed no
  tokens the gateway could report". A provider that generates output and only
  then drops the connection breaks that assumption, and because RunBudget.calls
  only increments inside charge(), the call ceiling never trips either - so an
  adversarial retry loop runs unbounded with the ledger reporting $0.00 while a
  real provider invoices $1.09. With the guard, 90 of 150 asks are refused
  pre-flight and the true bill is bounded by the attempt ceiling. Refusals are
  visible in exported telemetry (150 ERROR-status spans, 90 BudgetRefused
  nodes), both traces fetched back from Jaeger.

Also fixed, because it silently corrupted the headline number:
  s15code/economics/patient_transport.py. evals.yaml paces and retries the
  JUDGE path on the grounds that a rate limit is a transport failure and
  "giving up and calling the task unresolved would not be" honest - but that
  reasoning was never applied to the ANSWER path, where a 429 became an empty
  answer the judge then scored unresolved. A provider's per-minute quota was
  entering the resolution rate, the denominator of cost-per-resolved-task.
  Retries transient failures with backoff, raises permanent ones immediately,
  never fabricates a response. Live mode only; offline CI unchanged.

The one case the policy got wrong: on s12_two_issues_one_ticket the cascade
opened cheap, failed, escalated, and paid $0.00060875 where the frontier
baseline paid $0.00032675 - 86% more on that task.

tests/test_attempt_guard.py and tests/test_patient_transport.py cover both new
modules hermetically. No edit to the shipped controller.py or budget.py.
No secrets, no .env.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant