Add a code-review routing policy, its measurement, and an attack on it - #17
Open
nishanthvonteddu wants to merge 1 commit into
Open
nishanthvonteddu wants to merge 1 commit into
nishanthvonteddu wants to merge 1 commit into
Conversation
Session 15 assignment, Parts 1-3. Everything measured live against real
providers; no proof ran in offline mode.
The policy
A code-review workload (16 tasks, 4/5/7 by difficulty) with its own ladder
and budget policy in config_code_review/, selected with --config-dir so the
shipped config stays reproducible for Part 1.
economy cerebras / zai-glm-4.7 $0.00110 per call
standard groq / openai/gpt-oss-120b $0.00129
frontier gemini / gemini-3.1-flash-lite $0.00255
Every rung gets the same max_tokens (1600). The shipped ladder ramps 512 ->
4096, and a hard review needs up to 844 output tokens, so the cheap rung was
being cut off mid-answer and marked wrong for running out of room rather than
for being a weaker model. Equal ceilings isolate model quality; the cost is a
2.32x spread instead of 26x, which raises break-even from ~3.8% to 43.1%.
What the measurement found
A always-frontier 15/16 resolved $0.00061980 per resolved task
B always-cheapest 14/16 resolved $0.00042001
C budget-aware 16/16 resolved $0.00044066
The cascade resolves everything and is 28.9% cheaper per resolved task than
always-frontier. Against always-cheapest it LOSES by 4.7% -- but clears
break-even by only 1.8 points (87.5% vs 85.7%), so on 16 tasks the two are
indistinguishable on cost. What the cascade actually buys is the resolution
rate, not the saving. Reported as a tie rather than dressed up as a win.
t10_check_then_act: the frontier rung failed a race-condition review that both
cheaper rungs solved first try -- 1.46x the price for a worse answer.
t15_no_defect (a function with no bug, planted deliberately): the cheap model
invented a defect three times running. Retrying cannot fix confident
fabrication; changing model can.
The attack (proofs/p8_adversarial_code_review.py)
Aimed at the two controls this policy deliberately loosened. A runaway planner
attempted 120 rounds: 8 admitted, 112 refused, spend halted at 91.8% of the
ceiling. Before/after is measured, not extrapolated -- $0.01490625 uncontrolled
against $0.00917880 controlled over the same loop.
An unaffordable ceiling refused with 0 calls and $0 spent rather than
downgrading to something that still would not fit. One greedy node attempting
20 calls got exactly 4 (max_calls_per_node). Refused nodes emit no
provider_call span at all, and Jaeger carries the reason:
"BudgetRefused: ... spend pressure 0.918 >= refuse_at 0.9".
Limitations, stated rather than buried
- The judge cost $0.01712 against $0.02223 of work, and 35 samples were
unusable to rate limiting, so panel disagreement stopped being measurable
for those tasks. With the result resting on 1.8 points, that could flip it.
- Free tiers: nothing was billed. Costs are modelled from pricing.yaml.
- gemini-2.5-flash was removed as the frontier rung. glc_v4 guards its
thinking config with `if reasoning and reasoning != "off"`, so "off" sends
nothing and the model thinks anyway; its thought tokens then truncate the
answer AND are never metered, because the parser reads candidatesTokenCount
and ignores thoughtsTokenCount. Unmetered spend in a system whose invariant
is that no call escapes the ledger.
Also documents three commands in the session handout that silently do nothing:
GLC_OTEL_EXPORTER_ENDPOINT is read by no code, budget_usd is dropped by
RunBody so the run executes unbudgeted, and proofs without --base-url fall back
to a fake transport and pass with invented numbers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Session 15 assignment, Parts 1-3. A routing and budget policy for a code review
workload, measured against both baselines, then attacked.
Everything was run live against real providers. No proof ran in offline mode.
Where to read each part
docs/PART1_EVIDENCE.mddocs/PART2_PART3_EVIDENCE.md(first half)docs/PART2_PART3_EVIDENCE.md(second half) +proofs/p8_adversarial_code_review.pyREADME.mdStart with the README section if you want the short version; the two
docs/filesare the detailed working behind it.
Part 1 in one line
Both test suites pass (455 in
glc_v4, 277 here, higher than the handout's 445/244because these forks carry newer commits). All five proofs pass in live mode:
p19/9,p25/5,p36/6,p415/15,p710/10. Four runs captured with prompt,tier, model, event trace, Jaeger trace id, ledger rows and answer, including one
deliberately under-budgeted so a refusal appears in the evidence. The honest
limitation: the judge panel cost 16% more than every answer it graded, so any
saving has to be weighed against what it costs to know you achieved it.
The result
The cascade resolved every task and is 28.9% cheaper per resolved task than
always-frontier. Against always-cheapest it loses by 4.7% - but clears break-even
by only 1.8 percentage points (87.5% vs 85.7%), so on 16 tasks the two are
indistinguishable on cost. Reported as a tie rather than dressed up as a win. What
the cascade actually buys is the resolution rate: 100% vs 87.5%.
The ladder
Three providers, monotone cost, spread 2.32x, break-even 43.1%. Lives in
config_code_review/and is selected with--config-dir, so the shipped configstays reproducible for Part 1.
Every rung gets the same
max_tokens(1600). The shipped ladder ramps 512 -> 4096,and a hard review needs up to 844 output tokens - so the cheap rung was being cut off
mid-answer and marked wrong for running out of room rather than for being a weaker
model. Those two failures are indistinguishable in the final number. Equal ceilings
isolate model quality; the cost is a 2.32x spread instead of 26x, which raises
break-even from ~3.8% to 43.1%.
Where the policy is wrong
t10_check_then_act- the frontier rung failed a race-condition review thatboth cheaper rungs solved on the first attempt. 1.46x the price for a worse answer.
On this ladder price does not predict quality.
t15_no_defect- a function with no bug in it, planted deliberately. The cheapmodel invented a defect three times running. Retrying cannot fix confident
fabrication; changing model can, which is the one thing a cascade does that a retry
loop does not.
The attack (
proofs/p8_adversarial_code_review.py)Aimed at the two controls this policy deliberately loosened (
downgrade_at0.50->0.65,max_calls_per_node6->4). Before/after is measured, not extrapolated:The adversary never stopped asking. Spend halted at 91.8% of the ceiling because
refuse_at: 0.90said so. Two further attacks held: an unaffordable ceiling refusedwith 0 calls and $0 spent (rather than downgrading to something that still would
not fit), and one greedy node attempting 20 calls got exactly 4.
Refused nodes emit no
provider_callspan at all, and Jaeger carries the reason:Limitations, stated rather than buried
samples were unusable to rate limiting, so panel disagreement stopped being
measurable for those tasks. With the headline resting on 1.8 points, that could flip it.
pricing.yamlagainstreal reported token counts.
gemini-2.5-flashwas removed as the frontier rung.glc_v4guards its thinkingconfig with
if reasoning and reasoning != "off", so"off"sends nothing and themodel thinks anyway. Its thought tokens then truncate the answer and are never
metered, because the parser reads
candidatesTokenCountand ignoresthoughtsTokenCount. That is unmetered spend in a system whose stated invariant isthat no call escapes the ledger.
Three handout commands that silently do nothing
Each returns success while having no effect:
export GLC_OTEL_EXPORTER_ENDPOINT=...OTEL_EXPORTER_OTLP_ENDPOINT. The exporter no-ops when unconfigured, so traces vanish silently.{"budget_usd": 0.02}budget;RunBodydoes not forbid extras, so it is dropped and the run executes with no ceiling, returning HTTP 200.proofs/*.pywithout--base-urlContents
docs/PART1_EVIDENCE.mdproofs/tasks/code_review.py+.jsonlconfig_code_review/--config-dirdocs/PART2_PART3_EVIDENCE.mdproofs/p8_adversarial_code_review.pyREADME.mdconfig/tiers.yamlandconfig/pricing.yamlalso change: the shipped frontier rung(
github / openai/gpt-4.1) stopped answering when GitHub Models entered itsretirement brownout, so Part 1 could not run against the ladder as shipped. The
replacement and a rejected candidate are both documented in the file rather than
silently swapped.
Everything else is additive.
proofs/out/is gitignored, so the JSON results areregenerable output rather than part of this submission; the numbers in the docs are
the evidence of record.
No secrets,
.envfiles or user data are included. Verified: no key-shaped stringsin the diff, no real key value from any local
.envappears in any tracked file, anda fresh clone of this branch passes
uv run pytest -qat 277 passed.