Skip to content

A routing and budget policy for map/navigation reasoning, and what it cost - #18

Open
AshwaniBindroo-TomTom wants to merge 6 commits into
theschoolofai:mainfrom
ashwanibindroo-personal:s15-routing-policy
Open

AshwaniBindroo-TomTom wants to merge 6 commits into
theschoolofai:mainfrom
ashwanibindroo-personal:s15-routing-policy

Conversation

@AshwaniBindroo-TomTom

Copy link
Copy Markdown

A budget policy for an 18-task map/navigation reasoning workload, measured
against an always-frontier baseline, attacked, and reported with the places it
lost. Full evidence section at the end of the README; every number has a JSON
artefact under evidence/.

What this adds

Path What it is
proofs/tasks/navigation.jsonl the workload: 18 map/navigation reasoning tasks
config/navigation/ the policy: ladder, prices, budget thresholds, judge rubric
proofs/p0_calibration.py measures the ladder before any price is written down
proofs/p8_adversarial.py four attacks, spend before and refusal after
proofs/p9_run_capture.py one run captured whole: the six Part 1 artefacts
proofs/p10_judge_audit.py audits the judge against an answer key
proofs/p11_strategy_costs.py cost per resolved task, composed from measured rows
proofs/render_trace.py, show_refusals.py read traces back out of Jaeger
s15code/economics/ + tests/test_attempt_ceilings.py attempt ceilings, closing a hole p8 found

The measurement

strategy cost/call resolved cost/resolved
A always-frontier $0.00065911 14/18 $0.00084743
B always-cheapest $0.00046227 16/18 $0.00063563 (−25.0%)
C budget-aware $0.00041182 16/18 $0.00051478 (−39.3%)

Break-even for the cheap rung is 78.3%; it measured 88.9%, +10.6 points —
two tasks of margin.

The headline is not the saving. The cheap rung resolved more than the dear
one. Three of frontier's four failures were HTTP 429/503 on the call rather than
wrong answers, so per call that returned it resolved 14/15 against economy's
16/18 — an availability failure and a capability failure are indistinguishable in
a resolution rate, and only one is the model's fault.

Where the policy chose wrongly. The cascade escalated twice, paid
$0.00054650, and rescued neither task. Asking the cheap rung once and stopping
costs $0.00048063 per resolved task, so this policy is 7.1% worse than not
escalating at all
while beating always-frontier by 39%.

The attack that got through

A provider that generates an answer, burns the tokens and then fails returns
nothing to price, so it is never charged — and max_calls_per_* cannot see it,
because those count charges. Measured: 6 calls burned 252in/330out of real tokens
with the ledger recording 0 charges and 0 refusals. Attempt ceilings now
bound it (5 attempts, then a named refusal), capping exposure at $0.371940
instead of unbounded. Both defaults are 0, so no existing config changes.

Four defects found by running the documented instructions

  1. docker-compose.observability.yml pins jaegertracing/all-in-one:2.20.0; all-in-one is the v1 image and has no 2.x tags — v2 ships as jaegertracing/jaeger.
  2. glc/providers.py reads OPEN_ROUTER_API_KEY, documented nowhere; a correctly-named key loads and the provider is silently never built.
  3. The run instructions post budget_usd, but RunBody declares budget and does not set extra="forbid" — the documented call runs unbudgeted. Measured both ways in evidence/part3/.
  4. p1 scored a rate limit as an unresolved task, so throttling masqueraded as model failure, biased against whichever rung was under load. Fixed here behind opt-in flags (--answer-retries, --pace-seconds), defaults unchanged.

Honest scope

p1 did not finish: the Gemini free tier's daily quota ran out on every flash
model mid-study, and a single-provider ladder has no failover. Part 2 is composed
by p11 from p10's 36 measured, judged, priced rows — sequencing is arithmetic,
every cost and verdict is measured, and B's one assumption is verified live (3/3
byte-identical). Attack C could not reach a provider and its check fails saying
so. tests/ is 283 passing (277 upstream + 6 new); ruff clean.

🤖 Generated with Claude Code

…ned it

The task set is 18 map/navigation reasoning tasks, self-contained so they
measure arithmetic and rule interpretation rather than which gazetteer a model
memorised. Four of them are traps on purpose; without those the cheap rung
never fails and a ladder has nothing to show.

The ladder is two rungs because two is how many distinct prices this
deployment can actually reach. p0_calibration.py is why that sentence is a
measurement and not an assumption: it calls every rung and every judge before
any other proof is allowed to depend on them, and it asserts the four things a
configured ladder can quietly be lying about — a rung that 404s, a rung served
by a different model than its config names, a rung that returns "" at full
price, and an order that is not monotone in cost.

It has already earned its place twice. It caught gemini-3.1-pro-preview as
quota-blocked and gemini-3.1-flash as non-existent, which is what reduced a
planned three-rung ladder to an honest two. Reading its output then caught two
bugs in my own task file: nav05 asked for an 8-point compass sector but
expected WSW, which is a 16-point direction, and nav03 called 250 m at
motorway speed a gentle curve when it is 3.1 m/s^2 of lateral acceleration.

Measured, not assumed: spread 15.1x projected / 1.9x on the calibration
prompt, gemini-3.7-flash rejected as the top rung for a 25% HTTP 503 rate
against 12/12 for gemini-3.5-flash.
p8_adversarial.py runs four attacks: a runaway loop, a principal demanding a
rung it cannot pay for, a cascade driven up the whole ladder, and a provider
that generates an answer, consumes the tokens and THEN fails.

The fourth one gets through. The ledger charges from the token counts in a
response, so a call that dies after generation is unbilled real spend — and
max_calls_per_run/node cannot stop it either, because those count CHARGES and a
failed call never becomes one. A broken provider could be handed work forever
and no counter would move.

So the budget now counts attempts as well as charges. record_attempt() is
called before the transport is touched, max_attempts_per_run/node in
budgets.yaml bound it, and both default to 0 — disabled — so no existing
configuration changes behaviour. test_attempt_ceilings.py pins the fix and also
pins the BUG, so a future change that starts charging for failed calls has to
come and argue with a failing test rather than quietly closing the finding.
Exposure goes from unbounded to 30 x $0.012398 = $0.372, a number that fits in
a risk register.

Also here, because Part 1 asks for six artefacts per run and they live in six
different places: p9_run_capture.py records prompt, tier/model, ordered
journal events, Jaeger trace id, ledger rows and final answer for one run;
p10_judge_audit.py labels real answers twice, once mechanically from an answer
key and once by the panel, because every cost per resolved task divides by a
number two language models produced; render_trace.py and show_refusals.py read
traces back OUT of Jaeger, since a claim about telemetry should be checked
against what the collector holds rather than what the exporter believed it
sent.

render_trace.py summed every span carrying a cost and reported a run twice its
true price, because the run span carries the total and the provider-call spans
carry the parts. It now sums provider calls only and prints both figures with
whether they agree.
…tell me

Both suites, the five proofs against the shipped ladder offline and the
navigation ladder live, and four runs captured whole — prompt, tier and model,
ordered journal events, Jaeger trace id, ledger rows and final answer. The four
were chosen to cover different controller outcomes rather than four happy paths:
proceed, branch, refuse, proceed.

The honest limitation is one the traces exposed about themselves. Sorting run
1's spans by the start time Jaeger holds puts agent loop 2 at 0.0 ms and agent
loop 1 at 1.0 ms — loop 1 ran first, since it produced the plan that created
loop 2's node, so the trace states something causally impossible. The exporter
anchors a synthetic clock to the first metered call and lays untimed spans out
one millisecond apart in sequence order, so everything that happened before the
provider call is placed after it. No cost number is affected — the spans
reconcile with the ledger exactly, four times out of four — but these traces
cannot answer "where did the wall clock go", which is half of what people open
Jaeger for.

p7 fails two checks against this ladder and both failures are true: the rungs
are not spread across providers, and on one task the dearer rung answered in 28
output tokens where the cheap one spent 339, so it measurably cost LESS.
36 real answers, two rungs, labelled twice: mechanically from an answer key
that knows the right result, and by the panel with the rubric p1 uses. The
panel agreed with the key 36 times out of 36, with zero false resolves and zero
false unresolves — the two errors that would have pushed cost per resolved task
down and up respectively.

Worth stating plainly next to that: this is 36 answers on one task family with
crisp right answers, which is the easiest grading job there is, and the two
panel members agreed with each other on every single row — so this panel's
disagreement rate is unmeasured rather than low.

The audit also produced the structural finding for Part 2. The economy rung
resolves 16 of 18 and the frontier rung 14 of 18, so the cheap model resolves
MORE than the dear one. Three of frontier's four failures were HTTP 429 from
the gateway rather than wrong answers, which is a measurement contaminant and
is reported as one: excluding transport failures it is 14 of 15. Exactly one
task in eighteen is genuinely helped by escalating, and one is wrong at both
rungs.
evals.yaml already states the principle, for the judge: "a rate limit is a
TRANSPORT failure, not a verdict: retry it rather than let it become an
unresolved task." That reasoning was never applied to the answer path, and on a
rate-limited free tier the consequence is a row reading "unresolved,
cost=0.00000000" — indistinguishable in the aggregate from a model that
answered badly for nothing.

It does not merely add noise, it biases. The rung under the most sustained load
sheds the most rows, and cost per resolved task is computed from whatever
survives. The first live run of this workload lost its very first row that way,
with 23 of 71 calls on one key returning HTTP 429 while the daily quota sat at
74 of 1000 — the binding limit was per-MINUTE, and the proof had no way to say
so.

RetryingTransport wraps OUTSIDE MeteredTransport on purpose, so every attempt
still reaches the meter: `calls` stays the number that returned something to
charge for, `failures` the number that did not, and the retries are reported
next to them rather than hidden behind them. --answer-retries, --retry-backoff
and --pace-seconds all default to off, so a run that does not ask for them
behaves exactly as before.
Part 2. Cost per resolved task, against an always-frontier baseline:
always-frontier $0.00084743, always-cheapest $0.00063563, budget-aware
$0.00051478 — the cascade 39.3% cheaper per resolved task and 37.5% cheaper per
call. Break-even for the cheap rung is 78.3% from the measured 1.85x spread; it
sits at 88.9%, +10.6 points, which is two tasks of margin.

The headline is not the saving. The CHEAP rung resolved more tasks than the dear
one, 16 of 18 against 14 of 18 — price did not order these models by competence
at this workload. That is only mostly true and the qualification matters: three
of frontier's four failures were HTTP 429/503 on the call, not wrong answers, so
per call that RETURNED it resolved 14 of 15 against economy's 16 of 18. An
availability failure and a capability failure look identical in a resolution
rate and only one of them is the model's fault.

Where the policy chose wrongly: the cascade escalated on nav12 and nav14, paid
$0.00054650, and rescued neither — nav14 is wrong at both rungs and nav12's
frontier call never returned. Asking the cheap rung once and stopping costs
$0.00048063 per resolved task, so this budget-aware policy is 7.1% WORSE than
not escalating at all while beating always-frontier by 39%. A cascade bets the
rungs fail on different tasks; here their failures are nested.

p1 could not finish: the Gemini free tier's daily quota ran out on every flash
model mid-study, and a single-provider ladder has no failover — the cost of the
constraint tiers.yaml declares, arriving as a bill. p11 composes the three
strategies from p10's 36 measured, judged, priced rows, and MEASURES the one
assumption B rests on rather than asserting it: a temperature-0 retry returned a
byte-identical answer 3 times out of 3.

Part 3. Attack D got through and is the point: 6 calls burned 252in/330out of
real tokens with the ledger recording zero charges and zero refusals, because
charges need a response and the call ceilings count charges. Attempt ceilings
cut it off after 5 with a named refusal, bounding exposure at $0.371940 instead
of unbounded. A and B hold; C could not reach a provider and its check says so
rather than passing on a technicality. A fifth finding needs no attack at all:
the documented `budget_usd` field is not the one RunBody declares, and pydantic
drops it, so the session's own curl runs UNBUDGETED.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants