Skip to content

Add a code-review routing policy, its measurement, and an attack on it - #17

Open
nishanthvonteddu wants to merge 1 commit into
theschoolofai:mainfrom
nishanthvonteddu:s15-code-review-policy
Open

nishanthvonteddu wants to merge 1 commit into
theschoolofai:mainfrom
nishanthvonteddu:s15-code-review-policy

Conversation

@nishanthvonteddu

@nishanthvonteddu nishanthvonteddu commented Aug 12, 2026 •

Copy link
Copy Markdown

Session 15 assignment, Parts 1-3. A routing and budget policy for a code review
workload, measured against both baselines, then attacked.

Everything was run live against real providers. No proof ran in offline mode.

Where to read each part

Part What it covers Where
Part 1 - reproduce the floor Both suites green (455 + 277), all five shipped proofs passing live, four captured runs with Jaeger trace ids, and one honest limitation docs/PART1_EVIDENCE.md
Part 2 - build a policy and measure it The 16-task code-review set, the ladder, the budget policy, the judge, and the full comparison against both baselines docs/PART2_PART3_EVIDENCE.md (first half)
Part 3 - attack your own budget Three attacks on this policy's own controls, before/after spend, and the refusal in telemetry docs/PART2_PART3_EVIDENCE.md (second half) + proofs/p8_adversarial_code_review.py
Summary of all three The same evidence in plain language, with reproduction commands the Evidence section at the end of README.md

Start with the README section if you want the short version; the two docs/ files
are the detailed working behind it.

Part 1 in one line

Both test suites pass (455 in glc_v4, 277 here, higher than the handout's 445/244
because these forks carry newer commits). All five proofs pass in live mode:
p1 9/9, p2 5/5, p3 6/6, p4 15/15, p7 10/10. Four runs captured with prompt,
tier, model, event trace, Jaeger trace id, ledger rows and answer, including one
deliberately under-budgeted so a refusal appears in the evidence. The honest
limitation: the judge panel cost 16% more than every answer it graded, so any
saving has to be weighed against what it costs to know you achieved it.

The result

                    total spend  calls  cost per call  resolved  COST PER RESOLVED
A always-frontier    $0.0092970    16     $0.00058106   15/16       $0.00061980
B always-cheapest    $0.0058802    20     $0.00029401   14/16       $0.00042001
C budget-aware       $0.0070506    18     $0.00039170   16/16       $0.00044066

The cascade resolved every task and is 28.9% cheaper per resolved task than
always-frontier. Against always-cheapest it loses by 4.7% - but clears break-even
by only 1.8 percentage points (87.5% vs 85.7%), so on 16 tasks the two are
indistinguishable on cost. Reported as a tie rather than dressed up as a win. What
the cascade actually buys is the resolution rate: 100% vs 87.5%.

The ladder

economy   cerebras / zai-glm-4.7            $0.00110 per call
standard  groq / openai/gpt-oss-120b        $0.00129
frontier  gemini / gemini-3.1-flash-lite    $0.00255

Three providers, monotone cost, spread 2.32x, break-even 43.1%. Lives in
config_code_review/ and is selected with --config-dir, so the shipped config
stays reproducible for Part 1.

Every rung gets the same max_tokens (1600). The shipped ladder ramps 512 -> 4096,
and a hard review needs up to 844 output tokens - so the cheap rung was being cut off
mid-answer and marked wrong for running out of room rather than for being a weaker
model. Those two failures are indistinguishable in the final number. Equal ceilings
isolate model quality; the cost is a 2.32x spread instead of 26x, which raises
break-even from ~3.8% to 43.1%.

Where the policy is wrong

t10_check_then_act - the frontier rung failed a race-condition review that
both cheaper rungs solved on the first attempt. 1.46x the price for a worse answer.
On this ladder price does not predict quality.

t15_no_defect - a function with no bug in it, planted deliberately. The cheap
model invented a defect three times running. Retrying cannot fix confident
fabrication; changing model can, which is the one thing a cascade does that a retry
loop does not.

The attack (proofs/p8_adversarial_code_review.py)

Aimed at the two controls this policy deliberately loosened (downgrade_at 0.50->0.65,
max_calls_per_node 6->4). Before/after is measured, not extrapolated:

BEFORE  no effective ceiling    12 calls   $0.01490625     0 refusals
AFTER   the real policy          8 calls   $0.00917880   112 refusals
                                                         120 rounds attempted

The adversary never stopped asking. Spend halted at 91.8% of the ceiling because
refuse_at: 0.90 said so. Two further attacks held: an unaffordable ceiling refused
with 0 calls and $0 spent (rather than downgrading to something that still would
not fit), and one greedy node attempting 20 calls got exactly 4.

Refused nodes emit no provider_call span at all, and Jaeger carries the reason:

BudgetRefused: budget refused a frontier call for loop_9:
spend pressure 0.918 >= refuse_at 0.9

Limitations, stated rather than buried

  • The judge cost more than the work it graded ($0.01712 vs $0.02223), and 35
    samples were unusable to rate limiting, so panel disagreement stopped being
    measurable for those tasks. With the headline resting on 1.8 points, that could flip it.
  • Free tiers: nothing was billed. Costs are modelled from pricing.yaml against
    real reported token counts.
  • gemini-2.5-flash was removed as the frontier rung. glc_v4 guards its thinking
    config with if reasoning and reasoning != "off", so "off" sends nothing and the
    model thinks anyway. Its thought tokens then truncate the answer and are never
    metered, because the parser reads candidatesTokenCount and ignores
    thoughtsTokenCount. That is unmetered spend in a system whose stated invariant is
    that no call escapes the ledger.

Three handout commands that silently do nothing

Each returns success while having no effect:

Handout says Reality
export GLC_OTEL_EXPORTER_ENDPOINT=... read by no code; the real name is OTEL_EXPORTER_OTLP_ENDPOINT. The exporter no-ops when unconfigured, so traces vanish silently.
{"budget_usd": 0.02} the field is budget; RunBody does not forbid extras, so it is dropped and the run executes with no ceiling, returning HTTP 200.
proofs/*.py without --base-url proofs default to port 8112; the gateway is on 8111, and the harness silently substitutes a fake transport, passing with invented numbers.

Contents

File Part What it is
docs/PART1_EVIDENCE.md 1 Suites, five proofs, four captured runs, one honest limitation
proofs/tasks/code_review.py + .jsonl 2 The 16-task set (4 trivial / 5 moderate / 7 hard), generated from source so the embedded code is machine-escaped
config_code_review/ 2 The ladder, the budget policy and the relocated judge panel, selected with --config-dir
docs/PART2_PART3_EVIDENCE.md 2, 3 Full working for the measurement and the attack
proofs/p8_adversarial_code_review.py 3 The adversarial proof, 11 checks
README.md 1, 2, 3 The Evidence section: all three parts in plain language, plus reproduction commands

config/tiers.yaml and config/pricing.yaml also change: the shipped frontier rung
(github / openai/gpt-4.1) stopped answering when GitHub Models entered its
retirement brownout, so Part 1 could not run against the ladder as shipped. The
replacement and a rejected candidate are both documented in the file rather than
silently swapped.

Everything else is additive. proofs/out/ is gitignored, so the JSON results are
regenerable output rather than part of this submission; the numbers in the docs are
the evidence of record.

No secrets, .env files or user data are included. Verified: no key-shaped strings
in the diff, no real key value from any local .env appears in any tracked file, and
a fresh clone of this branch passes uv run pytest -q at 277 passed.

Session 15 assignment, Parts 1-3. Everything measured live against real
providers; no proof ran in offline mode.

The policy
  A code-review workload (16 tasks, 4/5/7 by difficulty) with its own ladder
  and budget policy in config_code_review/, selected with --config-dir so the
  shipped config stays reproducible for Part 1.

  economy   cerebras / zai-glm-4.7            $0.00110 per call
  standard  groq / openai/gpt-oss-120b        $0.00129
  frontier  gemini / gemini-3.1-flash-lite    $0.00255

  Every rung gets the same max_tokens (1600). The shipped ladder ramps 512 ->
  4096, and a hard review needs up to 844 output tokens, so the cheap rung was
  being cut off mid-answer and marked wrong for running out of room rather than
  for being a weaker model. Equal ceilings isolate model quality; the cost is a
  2.32x spread instead of 26x, which raises break-even from ~3.8% to 43.1%.

What the measurement found
  A always-frontier   15/16 resolved   $0.00061980 per resolved task
  B always-cheapest   14/16 resolved   $0.00042001
  C budget-aware      16/16 resolved   $0.00044066

  The cascade resolves everything and is 28.9% cheaper per resolved task than
  always-frontier. Against always-cheapest it LOSES by 4.7% -- but clears
  break-even by only 1.8 points (87.5% vs 85.7%), so on 16 tasks the two are
  indistinguishable on cost. What the cascade actually buys is the resolution
  rate, not the saving. Reported as a tie rather than dressed up as a win.

  t10_check_then_act: the frontier rung failed a race-condition review that both
  cheaper rungs solved first try -- 1.46x the price for a worse answer.

  t15_no_defect (a function with no bug, planted deliberately): the cheap model
  invented a defect three times running. Retrying cannot fix confident
  fabrication; changing model can.

The attack (proofs/p8_adversarial_code_review.py)
  Aimed at the two controls this policy deliberately loosened. A runaway planner
  attempted 120 rounds: 8 admitted, 112 refused, spend halted at 91.8% of the
  ceiling. Before/after is measured, not extrapolated -- $0.01490625 uncontrolled
  against $0.00917880 controlled over the same loop.

  An unaffordable ceiling refused with 0 calls and $0 spent rather than
  downgrading to something that still would not fit. One greedy node attempting
  20 calls got exactly 4 (max_calls_per_node). Refused nodes emit no
  provider_call span at all, and Jaeger carries the reason:
  "BudgetRefused: ... spend pressure 0.918 >= refuse_at 0.9".

Limitations, stated rather than buried
  - The judge cost $0.01712 against $0.02223 of work, and 35 samples were
    unusable to rate limiting, so panel disagreement stopped being measurable
    for those tasks. With the result resting on 1.8 points, that could flip it.
  - Free tiers: nothing was billed. Costs are modelled from pricing.yaml.
  - gemini-2.5-flash was removed as the frontier rung. glc_v4 guards its
    thinking config with `if reasoning and reasoning != "off"`, so "off" sends
    nothing and the model thinks anyway; its thought tokens then truncate the
    answer AND are never metered, because the parser reads candidatesTokenCount
    and ignores thoughtsTokenCount. Unmetered spend in a system whose invariant
    is that no call escapes the ledger.

Also documents three commands in the session handout that silently do nothing:
GLC_OTEL_EXPORTER_ENDPOINT is read by no code, budget_usd is dropped by
RunBody so the run executes unbudgeted, and proofs without --base-url fall back
to a fake transport and pass with invented numbers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant