Skip to content

S15: options-desk routing policy — one workload, two routing regimes, signature failure mode observed live - #15

Open
Shwethaamrutha wants to merge 3 commits into
theschoolofai:mainfrom
Shwethaamrutha:p15/feat/options-routing-policy
Open

Shwethaamrutha wants to merge 3 commits into
theschoolofai:mainfrom
Shwethaamrutha:p15/feat/options-routing-policy

Conversation

@Shwethaamrutha

@Shwethaamrutha Shwethaamrutha commented Aug 9, 2026 •

Copy link
Copy Markdown

A routing and budget policy for an Indian index-options trading desk, measured live against an always-frontier baseline and attacked. All numbers were generated by committed scripts against a live glc_v4 + Jaeger with real provider calls (Groq / Gemini / AWS Bedrock); the full evidence section is in the README, raw artefacts in evidence/.

The workload is deliberately two-faced

28 tasks (proofs/tasks/options_mixed.jsonl): 18 computational (spread breakevens, margin math, expiry logic — every answer Python-verified before the file was written) + 10 high-reasoning deliverables (RMS square-off compliance assessment, stop-loss-that-didn't-fire diagnosis, backtest code review with a planted look-ahead bug, expiry-day hold-or-close memo, the STT exercise trap, fund due-diligence, a production incident report…). The two halves get measurably different behaviour out of the same ladder — that contrast is the submission's core finding.

Headline (28 tasks, zero transport failures)

Strategy Cost/call Resolved Cost/resolved
A always-frontier (claude-sonnet-4-5 via Bedrock) $0.00972 28/28 $0.00972
B always-cheapest + 2 retries (gpt-oss-120b) $0.000295 24/28 $0.000466
C budget-aware cascade (this policy) $0.000356 28/28 $0.000407
  • signature failure mode OBSERVED (printed by p1 itself): B is 17.2% cheaper per call and 14.7% dearer per resolved task than C — the session's trap, sprung live, and only on the reasoning half.
  • C resolves 28/28 at 95.8% below the frontier baseline; its escalation gate fired on exactly the 4 tasks the cheap rung genuinely couldn't finish.
  • On the computational half alone, economy resolves 18/18 and the whole ladder is dead weight. Task type — not the price list — decides which regime you are in. Break-even arithmetic for both comparisons in §3.

Beyond the brief

  • A hidden-key correctness probe (probe_rungs.py): economy 18/18, standard 15/18, frontier 14/18 on the computational tasks — price does not order competence (§11, measured in my domain; frontier's only dropped p1 task was answered correctly by the cheapest model at 1/30th the price).
  • A truncation probe (probe_truncation.py): every failed economy attempt stopped at exactly the rung's 512-token cap; at 1500 the same model recovers r05 fully — part of what the ladder billed as capability was my own token cap (§6c).
  • Three real gateway defects found and fixed while getting honest numbers (shipped as gateway_setup/ patch): a client-side Groq TPM guard stricter than Groq's own live rate-limit headers; an instant-503 wait-loop for pinned providers that turned rolling-window limits into fake capability failures; and temperature-rejection healing for newer Bedrock Anthropic models (§2's 'heal the request', one more dialect).
  • The frontier rung bills real dollars (Bedrock, AWS credential chain, no keys through the gateway). A free top rung cannot be exhausted by a dollar ceiling — the shipped tiers.yaml makes this exact argument for the bottom rung — so refusals here are economically real: the adversarial run's 38 HTTP 402s each carry a real shortfall_usd.

Assignment checklist

  • Part 1: both suites green (glc_v4 455 passed / S15Code 277 passed), five proofs live, 4 captured runs with Jaeger trace IDs + ledger rows (capture_floor_runs.py), honest limitation: budget pressure silently served standard where the role requested frontier — visible only in the trace.
  • Part 2: ladder + budget policy on my own 28-task class; cost/call and cost/resolved vs always-frontier; break-even rates (5.9–8.6% vs A, ~89% vs C) with the workload placed against both; wrong cases in §6 (a,b,c) — each with its price.
  • Part 3: attack_wallet.py — 40 frontier-pinned rounds vs a $0.02 ceiling: 2 admitted, 38 refused, spend $0.0156, uncontrolled extrapolation $0.31, 38/38 refusals visible in /v1/refusals.
  • Reproduce-from-fresh-checkout commands in the README; no secrets, no .env, no user data.

Judge panel: Amazon Nova Pro + Llama 3.3 70B — disjoint from the answering ladder (asserted by p1). Judge meta-cost stated: $0.230 for 140 calls, 0.78× the entire answer spend.

…nature failure mode observed live

Part 2 policy + evidence for an Indian index-options trading desk:

- proofs/tasks/: 18 Python-verified computational tasks, 10 high-reasoning
  tasks (compliance assessment, incident report, code review, due diligence),
  the 28-task mixed set, plus escalation/extreme probe sets.
- config/: frontier rung repointed from retired GitHub Models gpt-4.1 to AWS
  Bedrock claude-sonnet-4-5; judge panel moved to Nova Pro + Llama 3.3 70B
  (disjoint from the answering ladder); Bedrock price rows.
- proofs/probe_rungs.py: per-rung correctness vs a hidden key — the capability
  inversion (economy 18/18, standard 15/18, frontier 14/18).
- proofs/attack_wallet.py: denial-of-wallet attack, 38/38 refusals recorded.
- proofs/capture_floor_runs.py: Part 1 evidence capture (trace IDs + ledger).
- evidence/: committed JSON from the live runs the README quotes.
- gateway_setup/: the glc_v4 patch (Bedrock provider, honest Groq TPM guard,
  window-wait for pinned providers, Opus temperature healing).

Headline (28 mixed tasks, zero transport failures): cascade resolves 28/28 at
95.8% below always-frontier; always-cheapest is 17.2% cheaper per call and
14.7% dearer per resolved task — the signature failure mode, observed live.
…ure (§6c)

Every economy attempt on the four escalated reasoning tasks stopped at exactly
the rung's 512-token cap (two with zero visible characters — the fully-billed
empty answer from §2 of the session). probe_truncation.py re-asks the same model
at a deliverable-sized cap: r04 completes cleanly and r05 recovers all required
content, so part of what the ladder billed as capability was the cap. The
insight that max_tokens is both the admission controller's worst case AND a
hidden failure mode is credited to the peer submission that right-sized caps
per task class; here it lands as a third honest wrong-case finding rather than
a re-measurement, so the published numbers stay true to the policy as measured.
…est limitation) and Part 3 log links for submission
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant