S15: an SRE-triage routing ladder measured on local models, and an attack on its budget - #21
Open
tanmays369 wants to merge 3 commits into
Open
tanmays369 wants to merge 3 commits into
tanmays369 wants to merge 3 commits into
Conversation
Add a reproducible local model ladder, adversarial budget proof, and honest evidence for task economics, routing failures, and Jaeger trace reconciliation.
Make the observability evidence directly reviewable while retaining its raw backend response for reconciliation.
Part 1 asks for the prompt and the ledger rows per run. Both were only reachable by opening docs/evidence JSON. Put them on the page, with the refusal record for the run that made no call at all.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Assignment 15. The workload is platform/SRE triage. Everything below was measured
live on this fork; no number is illustrative.
The honest constraint up front
This machine has no cloud provider credentials, so the ladder runs on two
independent local Ollama daemons and prices them against a declared rate
card in
config/pricing.yaml. Tokens, latencies, call counts, refusals,ledger rows and spans are all real and measured. The dollar amounts are
synthetic — they are the declared card applied to measured token counts. Every
ratio in this PR is therefore a consequence of that card, not of a provider
invoice. The README says this in its own words before it reports a single
figure, because a cost claim that hides its rate card is not evidence.
Making the ladder span more than one provider without cloud keys required a
companion change: an env-driven second local provider slot in the gateway,
tanmays369/glc_v4. It is a genuinely
separate daemon on its own port with its own queue, rate state and circuit
breaker — not a rename of the first slot, which is what
p7demands and what arename would have faked.
Part 1 — reproducing the floor
Both suites pass, and
p1,p2,p3,p4andp7run green. Four budgetedruns are captured with prompt, tier, model, ordered event trace, Jaeger trace
ID, ledger rows and final answer, including a Jaeger waterfall screenshot and
the backend JSON behind it.
The limitation the traces exposed: the refused run has nine agent spans and no
provider-call span, because no provider call ever existed. The refusal lives
in the graph journal and the budget
refusal_log, not as a billed model span.An operator whose dashboard watches provider spans would miss the system's
safest behaviour entirely. Refusals have to be ingested as a separate control
signal.
Part 2 — the policy, and what it cost
18 self-contained triage tasks in
proofs/tasks/sre_triage.jsonl(4 trivial, 6moderate, 8 hard). Ladder and role policy are data in
config/tiers.yaml:economy
llama3.2:1b, standardqwen2.5:3b, frontierqwen2.5:7bon thesecond slot. The controller downgrades at 50% pressure, refuses at 90%, and
reserves 20% for the terminal node.
Against always-frontier the policy cut cost per call by 76.25% and cost per
resolved task by 64.37%. The session's signature trap shows up cleanly between B
and C: B is 40.6% cheaper per call and 38.5% more expensive per resolved
task. Break-even against the frontier is r* = 18.5%; economy resolved 33.3%,
so it clears that bar — but against the cascade its break-even is 42.5%, and it
sits below it.
Judging is not free and is reported rather than buried: 46 judge calls cost
$0.01648380, which is 4.72x the policy workload it judged. That alone rules
this judge out of an online routing path. The judge (
phi4-mini) is disjointfrom every answering model, but it is a single local judge, so inter-judge
agreement cannot be measured — stated as a limitation rather than papered over.
Where the policy chose wrongly: on
s11_eks_api_tablethe frontier resolvedfor $0.00029000. The policy opened on economy, which put Deployment, Job and
Ingress on
v1; standard then invented software-style versions like2.5.0.Those two wrong attempts cost $0.00013320 and bought nothing. The frontier
escalation then hit a real 120s
ReadTimeout, so the task stayed unresolved.Two distinct causes — a cheap-first policy buying unusable answers, and a single
frontier slot with no failover — and the evidence records both rather than
blaming model quality.
Part 3 — attacking the budget
proofs/p8_attack_policy.pyruns four attacks and exits non-zero unless everycontrol holds.
80 rounds with 38 visible refusals. Loosened, every sampled round kept
spending; extrapolated to 10,000 rounds that is about $1.61 on the declared
card. Spend before the control, refusal after it.
Zero calls, zero spend, one refusal — the check happens before the call.
$0.00037440: a wrong confidence signal pays for every rung it passes.
two later errors reported no usage and were recorded as transport failures
rather than invented as charges.
Refusal visibility is the point: billed calls become ledger charges and provider
spans, while pre-call refusals become explicit graph failures and
refusal_logrecords. The two are never conflated.
Reviewing this
The four
config/*.yamlfiles repoint the shipped ladder at local models, sincea cloud ladder cannot run here.
S15_CONFIG_DIRlets a reviewer point theruntime back at any other config directory without touching code.
README.mdcarries the three evidence sections (
#part-1,#part-2,#part-3) and areproduce-from-fresh-checkout section with the exact daemon, gateway and proof
commands. Raw JSON for every proof and run is under
docs/evidence/.One test changed:
tests/test_cross_model_ladder.pyasserted that some pricedmodel carries the
measured_needs_reasoning_offannotation. Local models haveno thinking channel, so the assertion is now conditional on the annotation being
present — it still enforces the rule wherever the annotation exists, instead of
requiring the annotation to exist.