A routing and budget policy for map/navigation reasoning, and what it cost - #18
Open
AshwaniBindroo-TomTom wants to merge 6 commits into
Open
AshwaniBindroo-TomTom wants to merge 6 commits into
AshwaniBindroo-TomTom wants to merge 6 commits into
Conversation
…ned it The task set is 18 map/navigation reasoning tasks, self-contained so they measure arithmetic and rule interpretation rather than which gazetteer a model memorised. Four of them are traps on purpose; without those the cheap rung never fails and a ladder has nothing to show. The ladder is two rungs because two is how many distinct prices this deployment can actually reach. p0_calibration.py is why that sentence is a measurement and not an assumption: it calls every rung and every judge before any other proof is allowed to depend on them, and it asserts the four things a configured ladder can quietly be lying about — a rung that 404s, a rung served by a different model than its config names, a rung that returns "" at full price, and an order that is not monotone in cost. It has already earned its place twice. It caught gemini-3.1-pro-preview as quota-blocked and gemini-3.1-flash as non-existent, which is what reduced a planned three-rung ladder to an honest two. Reading its output then caught two bugs in my own task file: nav05 asked for an 8-point compass sector but expected WSW, which is a 16-point direction, and nav03 called 250 m at motorway speed a gentle curve when it is 3.1 m/s^2 of lateral acceleration. Measured, not assumed: spread 15.1x projected / 1.9x on the calibration prompt, gemini-3.7-flash rejected as the top rung for a 25% HTTP 503 rate against 12/12 for gemini-3.5-flash.
p8_adversarial.py runs four attacks: a runaway loop, a principal demanding a rung it cannot pay for, a cascade driven up the whole ladder, and a provider that generates an answer, consumes the tokens and THEN fails. The fourth one gets through. The ledger charges from the token counts in a response, so a call that dies after generation is unbilled real spend — and max_calls_per_run/node cannot stop it either, because those count CHARGES and a failed call never becomes one. A broken provider could be handed work forever and no counter would move. So the budget now counts attempts as well as charges. record_attempt() is called before the transport is touched, max_attempts_per_run/node in budgets.yaml bound it, and both default to 0 — disabled — so no existing configuration changes behaviour. test_attempt_ceilings.py pins the fix and also pins the BUG, so a future change that starts charging for failed calls has to come and argue with a failing test rather than quietly closing the finding. Exposure goes from unbounded to 30 x $0.012398 = $0.372, a number that fits in a risk register. Also here, because Part 1 asks for six artefacts per run and they live in six different places: p9_run_capture.py records prompt, tier/model, ordered journal events, Jaeger trace id, ledger rows and final answer for one run; p10_judge_audit.py labels real answers twice, once mechanically from an answer key and once by the panel, because every cost per resolved task divides by a number two language models produced; render_trace.py and show_refusals.py read traces back OUT of Jaeger, since a claim about telemetry should be checked against what the collector holds rather than what the exporter believed it sent. render_trace.py summed every span carrying a cost and reported a run twice its true price, because the run span carries the total and the provider-call spans carry the parts. It now sums provider calls only and prints both figures with whether they agree.
…tell me Both suites, the five proofs against the shipped ladder offline and the navigation ladder live, and four runs captured whole — prompt, tier and model, ordered journal events, Jaeger trace id, ledger rows and final answer. The four were chosen to cover different controller outcomes rather than four happy paths: proceed, branch, refuse, proceed. The honest limitation is one the traces exposed about themselves. Sorting run 1's spans by the start time Jaeger holds puts agent loop 2 at 0.0 ms and agent loop 1 at 1.0 ms — loop 1 ran first, since it produced the plan that created loop 2's node, so the trace states something causally impossible. The exporter anchors a synthetic clock to the first metered call and lays untimed spans out one millisecond apart in sequence order, so everything that happened before the provider call is placed after it. No cost number is affected — the spans reconcile with the ledger exactly, four times out of four — but these traces cannot answer "where did the wall clock go", which is half of what people open Jaeger for. p7 fails two checks against this ladder and both failures are true: the rungs are not spread across providers, and on one task the dearer rung answered in 28 output tokens where the cheap one spent 339, so it measurably cost LESS.
36 real answers, two rungs, labelled twice: mechanically from an answer key that knows the right result, and by the panel with the rubric p1 uses. The panel agreed with the key 36 times out of 36, with zero false resolves and zero false unresolves — the two errors that would have pushed cost per resolved task down and up respectively. Worth stating plainly next to that: this is 36 answers on one task family with crisp right answers, which is the easiest grading job there is, and the two panel members agreed with each other on every single row — so this panel's disagreement rate is unmeasured rather than low. The audit also produced the structural finding for Part 2. The economy rung resolves 16 of 18 and the frontier rung 14 of 18, so the cheap model resolves MORE than the dear one. Three of frontier's four failures were HTTP 429 from the gateway rather than wrong answers, which is a measurement contaminant and is reported as one: excluding transport failures it is 14 of 15. Exactly one task in eighteen is genuinely helped by escalating, and one is wrong at both rungs.
evals.yaml already states the principle, for the judge: "a rate limit is a TRANSPORT failure, not a verdict: retry it rather than let it become an unresolved task." That reasoning was never applied to the answer path, and on a rate-limited free tier the consequence is a row reading "unresolved, cost=0.00000000" — indistinguishable in the aggregate from a model that answered badly for nothing. It does not merely add noise, it biases. The rung under the most sustained load sheds the most rows, and cost per resolved task is computed from whatever survives. The first live run of this workload lost its very first row that way, with 23 of 71 calls on one key returning HTTP 429 while the daily quota sat at 74 of 1000 — the binding limit was per-MINUTE, and the proof had no way to say so. RetryingTransport wraps OUTSIDE MeteredTransport on purpose, so every attempt still reaches the meter: `calls` stays the number that returned something to charge for, `failures` the number that did not, and the retries are reported next to them rather than hidden behind them. --answer-retries, --retry-backoff and --pace-seconds all default to off, so a run that does not ask for them behaves exactly as before.
Part 2. Cost per resolved task, against an always-frontier baseline: always-frontier $0.00084743, always-cheapest $0.00063563, budget-aware $0.00051478 — the cascade 39.3% cheaper per resolved task and 37.5% cheaper per call. Break-even for the cheap rung is 78.3% from the measured 1.85x spread; it sits at 88.9%, +10.6 points, which is two tasks of margin. The headline is not the saving. The CHEAP rung resolved more tasks than the dear one, 16 of 18 against 14 of 18 — price did not order these models by competence at this workload. That is only mostly true and the qualification matters: three of frontier's four failures were HTTP 429/503 on the call, not wrong answers, so per call that RETURNED it resolved 14 of 15 against economy's 16 of 18. An availability failure and a capability failure look identical in a resolution rate and only one of them is the model's fault. Where the policy chose wrongly: the cascade escalated on nav12 and nav14, paid $0.00054650, and rescued neither — nav14 is wrong at both rungs and nav12's frontier call never returned. Asking the cheap rung once and stopping costs $0.00048063 per resolved task, so this budget-aware policy is 7.1% WORSE than not escalating at all while beating always-frontier by 39%. A cascade bets the rungs fail on different tasks; here their failures are nested. p1 could not finish: the Gemini free tier's daily quota ran out on every flash model mid-study, and a single-provider ladder has no failover — the cost of the constraint tiers.yaml declares, arriving as a bill. p11 composes the three strategies from p10's 36 measured, judged, priced rows, and MEASURES the one assumption B rests on rather than asserting it: a temperature-0 retry returned a byte-identical answer 3 times out of 3. Part 3. Attack D got through and is the point: 6 calls burned 252in/330out of real tokens with the ledger recording zero charges and zero refusals, because charges need a response and the call ceilings count charges. Attempt ceilings cut it off after 5 with a named refusal, bounding exposure at $0.371940 instead of unbounded. A and B hold; C could not reach a provider and its check says so rather than passing on a technicality. A fifth finding needs no attack at all: the documented `budget_usd` field is not the one RunBody declares, and pydantic drops it, so the session's own curl runs UNBUDGETED.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A budget policy for an 18-task map/navigation reasoning workload, measured
against an always-frontier baseline, attacked, and reported with the places it
lost. Full evidence section at the end of the README; every number has a JSON
artefact under
evidence/.What this adds
proofs/tasks/navigation.jsonlconfig/navigation/proofs/p0_calibration.pyproofs/p8_adversarial.pyproofs/p9_run_capture.pyproofs/p10_judge_audit.pyproofs/p11_strategy_costs.pyproofs/render_trace.py,show_refusals.pys15code/economics/+tests/test_attempt_ceilings.pyp8foundThe measurement
Break-even for the cheap rung is 78.3%; it measured 88.9%, +10.6 points —
two tasks of margin.
The headline is not the saving. The cheap rung resolved more than the dear
one. Three of frontier's four failures were HTTP 429/503 on the call rather than
wrong answers, so per call that returned it resolved 14/15 against economy's
16/18 — an availability failure and a capability failure are indistinguishable in
a resolution rate, and only one is the model's fault.
Where the policy chose wrongly. The cascade escalated twice, paid
$0.00054650, and rescued neither task. Asking the cheap rung once and stopping
costs $0.00048063 per resolved task, so this policy is 7.1% worse than not
escalating at all while beating always-frontier by 39%.
The attack that got through
A provider that generates an answer, burns the tokens and then fails returns
nothing to price, so it is never charged — and
max_calls_per_*cannot see it,because those count charges. Measured: 6 calls burned 252in/330out of real tokens
with the ledger recording 0 charges and 0 refusals. Attempt ceilings now
bound it (5 attempts, then a named refusal), capping exposure at $0.371940
instead of unbounded. Both defaults are
0, so no existing config changes.Four defects found by running the documented instructions
docker-compose.observability.ymlpinsjaegertracing/all-in-one:2.20.0;all-in-oneis the v1 image and has no 2.x tags — v2 ships asjaegertracing/jaeger.glc/providers.pyreadsOPEN_ROUTER_API_KEY, documented nowhere; a correctly-named key loads and the provider is silently never built.budget_usd, butRunBodydeclaresbudgetand does not setextra="forbid"— the documented call runs unbudgeted. Measured both ways inevidence/part3/.p1scored a rate limit as an unresolved task, so throttling masqueraded as model failure, biased against whichever rung was under load. Fixed here behind opt-in flags (--answer-retries,--pace-seconds), defaults unchanged.Honest scope
p1did not finish: the Gemini free tier's daily quota ran out on everyflashmodel mid-study, and a single-provider ladder has no failover. Part 2 is composed
by
p11fromp10's 36 measured, judged, priced rows — sequencing is arithmetic,every cost and verdict is measured, and B's one assumption is verified live (3/3
byte-identical). Attack C could not reach a provider and its check fails saying
so.
tests/is 283 passing (277 upstream + 6 new);ruffclean.🤖 Generated with Claude Code