Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
300 changes: 300 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,3 +210,303 @@ The generated protobuf modules under `s15code/core/a2a/` keep their original
filenames. They are reproduced verbatim because the serialized descriptor is
keyed on the `.proto` file name, and hand-editing generated gencode is worse than
a stale name.

---

# Evidence — a routing policy for code review, and what it really cost

Session 15 assignment. Everything below was run against real providers on
2026-08-11, not simulated. The short version:

- A cheap-to-strong cascade **resolved every task** (16/16) and was **29% cheaper
per resolved task** than always using the expensive model.
- But it **lost to always-using-the-cheapest** by 4.7%, by a margin so thin
(1.8 percentage points) that the honest answer is "these two are tied".
- The expensive model **failed a task both cheaper models solved**.
- An attack on the budget was refused 112 times out of 120, and the refusal is
visible in the trace with the reason attached.

Full working: [`docs/PART2_PART3_EVIDENCE.md`](docs/PART2_PART3_EVIDENCE.md).
Part 1 (reproducing the shipped floor): [`docs/PART1_EVIDENCE.md`](docs/PART1_EVIDENCE.md).

## The workload

**Code review.** Each task shows one function and asks for three things: name the
defect, describe a concrete scenario where it goes wrong, and give corrected code.

`proofs/tasks/code_review.jsonl` — 16 tasks, 4 trivial, 5 moderate, 7 hard.
Generated by `proofs/tasks/code_review.py` so the code snippets inside the JSON are
escaped by the machine rather than by hand.

Why the tasks demand long answers: an earlier run cost **5x** another *from the same
model* purely because it wrote more. If answers are short, every model costs about
the same and the whole comparison shows nothing.

One task, `t15_no_defect`, contains **no bug at all**. A model that invents one fails
it. That is the cheap model's characteristic failure — confident fabrication — and it
was planted on purpose so the policy had somewhere to be visibly wrong.

## The ladder

Three models, cheapest first. A "rung" is just a named model the policy can pick.

```
rung provider model price in/out cost per call (worst case)
economy cerebras zai-glm-4.7 0.50 / 0.50 $0.00110
standard groq openai/gpt-oss-120b 0.15 / 0.75 $0.00129
frontier gemini gemini-3.1-flash-lite 0.25 / 1.50 $0.00255
```

Three different providers, and cost rises at every step. Config:
`config_code_review/`, selected with `--config-dir` so the shipped config is
untouched.

**Every rung gets the same token limit (1600).** The shipped ladder gives the
expensive rung 8x more room than the cheap one. Measured here, a hard review needs
up to 844 tokens — so under the shipped 512-token cheap rung the answer gets cut off
mid-sentence and is marked wrong *for running out of room*, not for being a worse
model. Those two failures look identical in the final number. Equal limits mean the
only thing that differs between rungs is the model itself.

The price of that fairness: the gap between cheapest and dearest is **2.32x**, not
the shipped 26x. That makes the bar much harder to clear, which is the point.

## The numbers

Three strategies over the same 16 tasks, one ledger, two independent LLM judges
deciding whether each answer actually resolved its task.

```
total spend calls cost per call resolved COST PER RESOLVED
A always-frontier $0.0092970 16 $0.00058106 15/16 $0.00061980
B always-cheapest $0.0058802 20 $0.00029401 14/16 $0.00042001
C budget-aware $0.0070506 18 $0.00039170 16/16 $0.00044066
```

**Cost per resolved task is the number that matters.** Cost per call always flatters
the cheap model, because it ignores that a wrong answer has to be paid for twice.

Against the always-frontier baseline the cascade wins clearly: **28.9% cheaper per
resolved task, and it resolves more** (16/16 against 15/16).

## Break-even, and where we sit

Break-even is the resolution rate the cheap model must hit to be worth using. Below
it, retrying cheap costs more than just paying for the good model.

```
vs always-frontier break-even 73.0% cheap model got 87.5% -> 14.5 points clear
vs the cascade break-even 85.7% cheap model got 87.5% -> 1.8 points clear
```

The second line is the honest result. **The cascade lost.** Always-cheapest came out
4.7% cheaper per resolved task — but by 1.8 percentage points, on 16 tasks. Two tasks
either way would flip it.

So the conclusion is not "always-cheapest wins". It is that **at this ladder's 2.32x
spread the two strategies are indistinguishable on cost**, and what the cascade
actually buys is the resolution rate: 100% against 87.5%. If a missed review costs
more than $0.0002, the cascade is worth it. That is a business question, not a
measurement one, and the measurement says so rather than pretending otherwise.

## One case the policy got wrong

**`t10_check_then_act` — the most expensive model failed what both cheaper ones
solved.**

A race condition: two threads check "is this slot free?", both see yes, both write.
`gemini-3.1-flash-lite` (the frontier rung) did not resolve it. `zai-glm-4.7` and
`openai/gpt-oss-120b` both did, first attempt.

```
A spent $0.00058106 on the frontier rung and got nothing usable
C spent $0.00039790 on the cheapest rung and resolved it
```

Paying **1.46x for a worse answer**. On this ladder price does not predict quality —
which is why the policy opens on `standard` rather than `frontier`.

## The planted trap fired

```
t15_no_defect always-cheapest -> economy, economy, economy -> NEVER RESOLVED
budget-aware -> economy, then standard -> resolved
```

The function with no bug. The cheap model invented a defect three times running and
never backed down. Retrying the same model cannot fix confident fabrication;
**changing model can** — and that is the one thing a cascade does that a retry loop
does not. It is one of only two tasks the cheap-only strategy lost.

## A Jaeger trace with costs

Every run exports as a span tree. One provider call, with cost and tokens attached:

```
run
└─ agent_loop
└─ plan
└─ node
└─ provider_call gemini-2.5-flash
gen_ai.usage.input_tokens 227
gen_ai.usage.output_tokens 503
s15.cost 0.0013256
```

The trace and the ledger are built from the same journal by different readers, and
`p4` checks they agree: **span costs summed to the ledger with a difference of
exactly 0.000e+00**. No prompt or answer text reaches the backend — verified against
the live backend, not just asserted in config.

## The attack

`proofs/p8_adversarial_code_review.py` — written against *this* policy, aimed at the
two controls it deliberately loosened (`downgrade_at` 0.50→0.65,
`max_calls_per_node` 6→4). All 11 checks pass.

The adversary is a planner that earns a new task from every outcome, forever, always
asking for the most expensive rung. Run twice — once with the ceiling raised so high
it never bites, once with the real policy:

```
BEFORE no effective ceiling 12 calls spent $0.01490625 0 refusals
AFTER the real policy 8 calls spent $0.00917880 112 refusals
120 rounds attempted
```

It never stopped asking. **120 attempts, 112 refused, 8 allowed.** Spend stopped at
91.8% of the ceiling because `refuse_at: 0.90` said so — not because the attacker
gave up. Extrapolated from this run's own measured cost per call, the uncontrolled
version bills **~$12.42 over 10,000 rounds**.

Two more targeted attacks, both held:

```
unaffordable ceiling asked for the dearest rung with less money than the CHEAPEST
rung costs -> 0 calls, $0 spent, 1 refusal
one greedy node a single task calling the model 20 times
-> 4 allowed, then refused (max_calls_per_node = 4, exactly)
```

The first is the subtle one: it **refused rather than downgrading** to something that
still would not fit, and contacted no provider at all.

## The refusal is visible in telemetry

Fetched back out of Jaeger by trace id, the refused nodes carry the reason:

```
otel.status_description
BudgetRefused: budget refused a frontier call for loop_9:
spend pressure 0.918 >= refuse_at 0.9
```

The control, the threshold, and the measured pressure that tripped it. A trace saying
only "this failed" cannot audit a budget.

And **refused nodes emit no `provider_call` span at all** — 8 provider-call spans for
8 allowed calls, against 120 nodes. The absence is the evidence that nothing was sent.

## Honest limitations

**1. Judging cost more than the work, and the panel degraded.**

```
judge 37 calls $0.01711875 35 unusable samples 226 transport retries
work $0.02223
```

The judge was heavily rate-limited, so many verdicts fell back to a single grader
instead of a two-member panel. Verdicts still parsed, but *disagreement between
judges* stopped being measurable. With the headline result resting on 1.8 percentage
points, this is a real threat to the conclusion, not a footnote. A re-run with a
rate-limit-tolerant panel could flip it.

**2. Nothing was actually billed.** All providers are on free tiers. Every dollar
figure is computed from `pricing.yaml` rates against real reported token counts. The
control logic is real; the invoice is not.

**3. A gateway defect makes one model's spend invisible.** `gemini-2.5-flash` was the
frontier rung for one run and had to be removed. `glc_v4/glc/providers.py` guards its
thinking config with `if reasoning and reasoning != "off"`, so asking for
`reasoning: "off"` sends nothing — and that model thinks by default, so "off" leaves
thinking **on**. Its thought tokens then eat the answer budget (8 of 12 calls returned
~61 visible tokens, truncated mid-review) and are **never metered**, because the
parser reads `candidatesTokenCount` and ignores `thoughtsTokenCount`. Google bills
them; this ledger never sees them. In a system whose stated invariant is that no call
escapes the ledger, that is unmetered spend.

**4. Downgrading is a weak control on a narrow ladder.** At 2.32x, dropping a rung
saves 57%; on the shipped 26x ladder it saves 96%. This policy therefore leans on the
absolute ceiling and the call counters rather than on graceful degradation.

## Reproducing this from a fresh checkout

```bash
git clone https://github.com/theschoolofai/glc_v4.git
git clone https://github.com/theschoolofai/S15Code.git

# Jaeger. No Docker needed — this is the project's own fallback.
glc_v4/scripts/jaeger_local.sh --start

# Gateway. Keys go in glc_v4/.env and are never committed.
cd glc_v4 && uv sync && uv run glc serve # port 8111

# Runtime, in a second terminal.
cd S15Code && uv sync && uv run s15code serve # port 8113

cd S15Code
B=http://127.0.0.1:8111
OTLP=http://127.0.0.1:4318/v1/traces

python proofs/tasks/code_review.py # regenerate the task set

# Part 2 — the measurement. Takes about 2.4 hours; the judge is rate-limited.
uv run python proofs/p1_cost_per_task.py \
--tasks proofs/tasks/code_review.jsonl \
--config-dir config_code_review \
--principal course/s15/part2 --label code_review_v2 --base-url $B

# Part 3 — the attack.
uv run python proofs/p8_adversarial_code_review.py \
--task "adversarial budget test" \
--config-dir config_code_review --budget 0.01 --uncontrolled-rounds 12 \
--principal course/s15/part3 --label code_review \
--base-url $B --otel-endpoint $OTLP
```

**`--base-url` is required.** The proofs default to port 8112, and when no gateway
answers there they silently swap in a fake transport and pass anyway — with invented
numbers and no trace. Two other commands in the session handout are also silently
wrong; both are documented in `docs/PART1_EVIDENCE.md`.

The environment variable for tracing is `OTEL_EXPORTER_OTLP_ENDPOINT`, not
`GLC_OTEL_EXPORTER_ENDPOINT`. Nothing reads the latter, and the exporter is a no-op
when unconfigured, so the gateway looks healthy and exports nothing.

**Running the test suite.** On a clean checkout it is simply green:

```bash
uv sync && uv run pytest -q # 277 passed
```

Verified by cloning this branch into an empty directory and running it there.

It only goes red once you add a `.env` configuring a real collector for the proofs.
`s15code/main.py` calls `load_dotenv`, so the tests read your local `.env`, and
`test_the_trace_route_serves_the_span_tree` asserts `exported_over_the_wire is False`
on the grounds that no collector is needed. True on a clean checkout; false the
moment you point the runtime at Jaeger.

The fix is to set the variable **empty** rather than to unset it:

```bash
S15_OTEL_EXPORTER_ENDPOINT= uv run pytest -q # 277 passed
env -u S15_OTEL_EXPORTER_ENDPOINT uv run pytest -q # still fails
```

`load_dotenv` does not override a variable that is already set, so an empty value
wins, while unsetting it just lets the `.env` value back in. The suite is not
hermetic with respect to the developer's environment.

Results land in `proofs/out/`, which is gitignored — the numbers above are the record.
9 changes: 9 additions & 0 deletions config/pricing.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,15 @@ models:
measured_non_empty: true
gemini-3.1-flash: {input: 0.50, output: 3.00}
gemini-3.1-pro: {input: 2.00, output: 12.00}
# LADDER rung 3 (frontier) as of 2026-08-11. Rates are NOT invented here: they
# are copied from glc_v4's own price table (glc/economics/pricing.yaml), which
# is the table the gateway actually bills this model against — so the ledger and
# this file cannot disagree. MEASURED through the gateway: 8 in / 2 out,
# $0.0000049, price_source "model".
#
# Why this model and not gemini-3.1-pro: the free AI Studio tier has no quota for
# pro (HTTP 429 on every attempt), and a rung that cannot be called is not a rung.
gemini-2.5-flash: {input: 0.30, output: 2.50}
# Open weight / hosted (the other models this gateway serves today)
#
# NOT REACHABLE as of 2026-07-30: this is the model NVIDIA_MODEL selects, and
Expand Down
Loading