S15: options-desk routing policy — one workload, two routing regimes, signature failure mode observed live - #15
Open
Shwethaamrutha wants to merge 3 commits into
Conversation
…nature failure mode observed live Part 2 policy + evidence for an Indian index-options trading desk: - proofs/tasks/: 18 Python-verified computational tasks, 10 high-reasoning tasks (compliance assessment, incident report, code review, due diligence), the 28-task mixed set, plus escalation/extreme probe sets. - config/: frontier rung repointed from retired GitHub Models gpt-4.1 to AWS Bedrock claude-sonnet-4-5; judge panel moved to Nova Pro + Llama 3.3 70B (disjoint from the answering ladder); Bedrock price rows. - proofs/probe_rungs.py: per-rung correctness vs a hidden key — the capability inversion (economy 18/18, standard 15/18, frontier 14/18). - proofs/attack_wallet.py: denial-of-wallet attack, 38/38 refusals recorded. - proofs/capture_floor_runs.py: Part 1 evidence capture (trace IDs + ledger). - evidence/: committed JSON from the live runs the README quotes. - gateway_setup/: the glc_v4 patch (Bedrock provider, honest Groq TPM guard, window-wait for pinned providers, Opus temperature healing). Headline (28 mixed tasks, zero transport failures): cascade resolves 28/28 at 95.8% below always-frontier; always-cheapest is 17.2% cheaper per call and 14.7% dearer per resolved task — the signature failure mode, observed live.
…ure (§6c) Every economy attempt on the four escalated reasoning tasks stopped at exactly the rung's 512-token cap (two with zero visible characters — the fully-billed empty answer from §2 of the session). probe_truncation.py re-asks the same model at a deliverable-sized cap: r04 completes cleanly and r05 recovers all required content, so part of what the ladder billed as capability was the cap. The insight that max_tokens is both the admission controller's worst case AND a hidden failure mode is credited to the peer submission that right-sized caps per task class; here it lands as a third honest wrong-case finding rather than a re-measurement, so the published numbers stay true to the policy as measured.
…est limitation) and Part 3 log links for submission
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A routing and budget policy for an Indian index-options trading desk, measured live against an always-frontier baseline and attacked. All numbers were generated by committed scripts against a live glc_v4 + Jaeger with real provider calls (Groq / Gemini / AWS Bedrock); the full evidence section is in the README, raw artefacts in
evidence/.The workload is deliberately two-faced
28 tasks (
proofs/tasks/options_mixed.jsonl): 18 computational (spread breakevens, margin math, expiry logic — every answer Python-verified before the file was written) + 10 high-reasoning deliverables (RMS square-off compliance assessment, stop-loss-that-didn't-fire diagnosis, backtest code review with a planted look-ahead bug, expiry-day hold-or-close memo, the STT exercise trap, fund due-diligence, a production incident report…). The two halves get measurably different behaviour out of the same ladder — that contrast is the submission's core finding.Headline (28 tasks, zero transport failures)
claude-sonnet-4-5via Bedrock)gpt-oss-120b)signature failure mode OBSERVED(printed by p1 itself): B is 17.2% cheaper per call and 14.7% dearer per resolved task than C — the session's trap, sprung live, and only on the reasoning half.Beyond the brief
probe_rungs.py): economy 18/18, standard 15/18, frontier 14/18 on the computational tasks — price does not order competence (§11, measured in my domain; frontier's only dropped p1 task was answered correctly by the cheapest model at 1/30th the price).probe_truncation.py): every failed economy attempt stopped at exactly the rung's 512-token cap; at 1500 the same model recovers r05 fully — part of what the ladder billed as capability was my own token cap (§6c).gateway_setup/patch): a client-side Groq TPM guard stricter than Groq's own live rate-limit headers; an instant-503 wait-loop for pinned providers that turned rolling-window limits into fake capability failures; andtemperature-rejection healing for newer Bedrock Anthropic models (§2's 'heal the request', one more dialect).shortfall_usd.Assignment checklist
capture_floor_runs.py), honest limitation: budget pressure silently servedstandardwhere the role requestedfrontier— visible only in the trace.attack_wallet.py— 40 frontier-pinned rounds vs a $0.02 ceiling: 2 admitted, 38 refused, spend $0.0156, uncontrolled extrapolation $0.31, 38/38 refusals visible in/v1/refusals..env, no user data.Judge panel: Amazon Nova Pro + Llama 3.3 70B — disjoint from the answering ladder (asserted by p1). Judge meta-cost stated: $0.230 for 140 calls, 0.78× the entire answer spend.