Skip to content

S15: an SRE-triage routing ladder measured on local models, and an attack on its budget - #21

Open
tanmays369 wants to merge 3 commits into
theschoolofai:mainfrom
tanmays369:s15-sre-triage-routing-policy
Open

tanmays369 wants to merge 3 commits into
theschoolofai:mainfrom
tanmays369:s15-sre-triage-routing-policy

Conversation

@tanmays369

Copy link
Copy Markdown

Assignment 15. The workload is platform/SRE triage. Everything below was measured
live on this fork; no number is illustrative.

The honest constraint up front

This machine has no cloud provider credentials, so the ladder runs on two
independent local Ollama daemons and prices them against a declared rate
card
in config/pricing.yaml. Tokens, latencies, call counts, refusals,
ledger rows and spans are all real and measured. The dollar amounts are
synthetic
— they are the declared card applied to measured token counts. Every
ratio in this PR is therefore a consequence of that card, not of a provider
invoice. The README says this in its own words before it reports a single
figure, because a cost claim that hides its rate card is not evidence.

Making the ladder span more than one provider without cloud keys required a
companion change: an env-driven second local provider slot in the gateway,
tanmays369/glc_v4. It is a genuinely
separate daemon on its own port with its own queue, rate state and circuit
breaker — not a rename of the first slot, which is what p7 demands and what a
rename would have faked.

Part 1 — reproducing the floor

Both suites pass, and p1, p2, p3, p4 and p7 run green. Four budgeted
runs are captured with prompt, tier, model, ordered event trace, Jaeger trace
ID, ledger rows and final answer, including a Jaeger waterfall screenshot and
the backend JSON behind it.

The limitation the traces exposed: the refused run has nine agent spans and no
provider-call span
, because no provider call ever existed. The refusal lives
in the graph journal and the budget refusal_log, not as a billed model span.
An operator whose dashboard watches provider spans would miss the system's
safest behaviour entirely. Refusals have to be ingested as a separate control
signal.

Part 2 — the policy, and what it cost

18 self-contained triage tasks in proofs/tasks/sre_triage.jsonl (4 trivial, 6
moderate, 8 hard). Ladder and role policy are data in config/tiers.yaml:
economy llama3.2:1b, standard qwen2.5:3b, frontier qwen2.5:7b on the
second slot. The controller downgrades at 50% pressure, refuses at 90%, and
reserves 20% for the terminal node.

Strategy Spend Calls Cost/call Resolved Cost/resolved task
A: always frontier $0.00882300 18 $0.00049017 9/18 $0.00098033
B: always economy, two retries $0.00290220 42 $0.00006910 6/18 $0.00048370
C: budget-aware cascade $0.00349260 30 $0.00011642 10/18 $0.00034926

Against always-frontier the policy cut cost per call by 76.25% and cost per
resolved task by 64.37%. The session's signature trap shows up cleanly between B
and C: B is 40.6% cheaper per call and 38.5% more expensive per resolved
task.
Break-even against the frontier is r* = 18.5%; economy resolved 33.3%,
so it clears that bar — but against the cascade its break-even is 42.5%, and it
sits below it.

Judging is not free and is reported rather than buried: 46 judge calls cost
$0.01648380, which is 4.72x the policy workload it judged. That alone rules
this judge out of an online routing path. The judge (phi4-mini) is disjoint
from every answering model, but it is a single local judge, so inter-judge
agreement cannot be measured — stated as a limitation rather than papered over.

Where the policy chose wrongly: on s11_eks_api_table the frontier resolved
for $0.00029000. The policy opened on economy, which put Deployment, Job and
Ingress on v1; standard then invented software-style versions like 2.5.0.
Those two wrong attempts cost $0.00013320 and bought nothing. The frontier
escalation then hit a real 120s ReadTimeout, so the task stayed unresolved.
Two distinct causes — a cheap-first policy buying unusable answers, and a single
frontier slot with no failover — and the evidence records both rather than
blaming model quality.

Part 3 — attacking the budget

proofs/p8_attack_policy.py runs four attacks and exits non-zero unless every
control holds.

  • Runaway loop. Under a $0.002 ceiling spend flattened at $0.001764 across
    80 rounds with 38 visible refusals. Loosened, every sampled round kept
    spending; extrapolated to 10,000 rounds that is about $1.61 on the declared
    card. Spend before the control, refusal after it.
  • Unaffordable tier. Projected $0.00021060 against $0.00000029 remaining.
    Zero calls, zero spend, one refusal — the check happens before the call.
  • Forced full cascade. The same task climbed all three rungs for
    $0.00037440: a wrong confidence signal pays for every rung it passes.
  • Provider failure after token consumption. The successful call was charged;
    two later errors reported no usage and were recorded as transport failures
    rather than invented as charges.

Refusal visibility is the point: billed calls become ledger charges and provider
spans, while pre-call refusals become explicit graph failures and refusal_log
records. The two are never conflated.

Reviewing this

The four config/*.yaml files repoint the shipped ladder at local models, since
a cloud ladder cannot run here. S15_CONFIG_DIR lets a reviewer point the
runtime back at any other config directory without touching code. README.md
carries the three evidence sections (#part-1, #part-2, #part-3) and a
reproduce-from-fresh-checkout section with the exact daemon, gateway and proof
commands. Raw JSON for every proof and run is under docs/evidence/.

One test changed: tests/test_cross_model_ladder.py asserted that some priced
model carries the measured_needs_reasoning_off annotation. Local models have
no thinking channel, so the assertion is now conditional on the annotation being
present — it still enforces the rule wherever the annotation exists, instead of
requiring the annotation to exist.

Add a reproducible local model ladder, adversarial budget proof, and honest evidence for task economics, routing failures, and Jaeger trace reconciliation.
Make the observability evidence directly reviewable while retaining its raw backend response for reconciliation.
Part 1 asks for the prompt and the ledger rows per run. Both were only
reachable by opening docs/evidence JSON. Put them on the page, with the
refusal record for the run that made no call at all.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant