Skip to content

feat: budget-aware routing policy for LLM cost engineering, measured live - #7

Open
mkthoma wants to merge 8 commits into
theschoolofai:mainfrom
mkthoma:feat/cost-engineering-policy-and-evidence
Open

mkthoma wants to merge 8 commits into
theschoolofai:mainfrom
mkthoma:feat/cost-engineering-policy-and-evidence

Conversation

@mkthoma

@mkthoma mkthoma commented Aug 7, 2026 •

Copy link
Copy Markdown

A budget-aware routing policy for a LLM / cloud cost engineering workload, with the measurement, the break-even analysis, and the places it goes wrong. All numbers measured live 2026-08-07 against glc_v4.

Jump links are in a box at the top of README.md. The evidence sections are
Part 1, then sections 1-7 for Parts 2 and 3; full working in docs/, and the
raw live results are committed under docs/evidence/.


The ladder had to be rebuilt first

config/tiers.yaml records its three rungs as measured on 2026-07-30. GitHub Models - which served the frontier rung, openai/gpt-4.1 - was fully retired on 2026-07-30, the same day.

Verified dead on 2026-08-06: 17 model IDs across 7 vendors all return HTTP 410 github_models_retirement_brownout, as does GET /catalog/models and a request naming no model at all. A deliberately invalid ID returns 410 rather than 404, proving the endpoint short-circuits before model resolution.

rung provider model $/Mtok in/out
economy groq openai/gpt-oss-120b 0.15 / 0.75
standard gemini gemini-3.5-flash-lite 0.30 / 2.50
frontier nvidia nvidia/nemotron-3-ultra-550b-a55b 0.60 / 3.60

Frontier priced from OpenRouter's published rate for the identical model, since NVIDIA's free tier bills $0 and a free top rung makes "refuse at exhaustion" inexpressible in dollars.

Measured, against an always-frontier baseline

Two arms, identical but for the token ceiling - same tasks, models, prices, judge.

Wide ladder (economy ceiling 512):

strategy spend calls cost/call resolved cost/resolved
A always-frontier (baseline) 0.00509460 17 0.00029968 10/17 0.00050946
B always-cheapest 0.00914595 27 0.00033874 14/17 0.00065328
C budget-aware 0.01086335 21 0.00051730 17/17 0.00063902

The signature failure mode was observed (B vs C): cost per call -34.5% while cost per resolved task rose +2.2%. Change the denominator, reverse the ranking.

The measured ladder is inverted. Projected 32.5x with frontier dearest; measured, frontier costs less per call than economy - the economy model writes 7.3x more output tokens (407.6 vs 55.6 mean), which more than cancels frontier's 4.8x output rate.

Break-even

r* = k / (C/c - 1 + k), k = 3. In the wide arm the cheap rung sat 3.3 points below break-even (0.824 against 0.8562), which is why it lost.

What moved it was not price - it was truncation. Raising the economy ceiling 512->1024 cut truncated attempts from 18 to 4 and lifted resolution 82.4% -> 94.1%, carrying it well above break-even. Strategy A is the control at 10/17 in both arms, since it writes 5-24 output tokens and never nears a ceiling.

My original hypothesis - that a narrower price spread would make the trap appear - was wrong, and the confound is stated plainly in the docs: equalising max_tokens changed the ceiling and the projected spread together, because the projection is a function of the ceiling.

Where the policy chose wrongly

c06_cache_discount:

B economy 251 out $0.00022155 RESOLVED overall 1.0
C economy 512 out $0.00041730 FAILED overall 0.5 <- truncated at the cap
C standard 432 out $0.00112770 RESOLVED overall 1.0
   C total $0.00154500 = 7.0x what B spent, for the same final answer

Same model, same prompt, temperature 0 - one run wrote 251 tokens and finished, the other ran to exactly 512 and was cut off. The cascade cannot distinguish "the model was wrong" from "the model ran out of room."

Reported beside c11, where the same rule climbed all three rungs at 2.8x and reached an answer B never did.

Capability is not monotonic. The frontier rung resolved 1 of 5 moderate tasks where economy resolved 5 of 5, failing with arithmetic slips in 5-24 output tokens (0.018360 for 0.018346; $3.30 for $4.30). The cause is this policy's own reasoning: "off" - set to avoid the invisible-token billing bug, it also removed the model's scratchpad. The dial that stops you paying for invisible tokens is the same dial that costs you correctness on multi-step work.

Every frontier miss scored exactly 0.6667 - 4 on all four generic criteria, 0 on meets_expectation. Precisely the case the 0.5 per-criterion floor exists to catch.

The adversarial test

proofs/tasks/verbosity_attack.jsonl - six requests whose correct answers are long, paired with six controls from the same domain that are equally easy but short. Every request is individually cheap and passes any spend ceiling.

group n mean/request vs control
control 6 $0.00004825 1.00x
attack, no escalation 4 $0.00035404 7.34x - verbosity alone
attack, escalated 2 $0.00228495 47.36x - verbosity + ladder climb

20.7x amplification per request, all 12 resolved - it inflates cost while looking like a completely successful workload. The escalation multiplier of 6.45x independently reproduces the 7.0x measured on c06. One request missed the cap by six tokens; asking for "integers 1 to 400" makes it deterministic.

Before/after the control (p3, 60 rounds, $0.0005): 1 admitted, 59 refused, 59 nodes failed with BudgetRefused.

Limitations found while reproducing the floor

docs/LIMITATIONS.md - ten trace-derived, nine environmental, each marked OBSERVED or INFERRED. The ones that may matter to this repo:

  • p1 silently simulates and reports success. It defaults to port 8112 while every other proof and GLC_PORT use 8111, and no proof loads .env, so GLC_BASE_URL is never read on that path. A live-intended run against a healthy gateway produced a full cost table, a break-even calculation and a signature failure mode OBSERVED finding - all simulated, exit 0, ok: true, every gate green.
  • p4's backend gates race ingestion - 8 of 10 spans on one run, 7 of 10 on the next; re-querying seconds later returns all 10. No wait or retry between export and assertion.
  • p4's query URL is unreachable on Windows/WSL - Jaeger binds [::] and WSL will not relay that to the IPv4 literal. S15_TRACE_QUERY_URL=http://localhost:16686 works.
  • A run has no stable trace id - /v1/agent/runs/{id}/trace re-exports and mints a new one per call; Jaeger held 11 traces for 5 runs, one run in 5 copies.
  • An unbudgeted run spends real money and reports none of it - no controller in the path means no charge recorded: provider_calls=0, cost=0 on a run that made a real billable call.
  • GEMINI_MODEL defaults to the retired gemini-2.5-flash, which fails every unbudgeted run.
  • jaeger_local.sh rejects Windows despite Jaeger shipping a Windows binary, and ships CRLF that breaks it under WSL.
  • Two shipped tests fail on a clean checkout (one is a POSIX file:// assumption).

What's in the diff

path what
proofs/tasks/llm_cost_engineering.jsonl 17 tasks, 4/5/8, expectations independently recomputed
proofs/tasks/verbosity_attack.jsonl 6 attack + 6 control, the adversarial set
config/tiers.yaml, config/pricing.yaml ladder repointed, prices sourced and dated
config_narrow/ the crossover control arm
docs/*.md full evidence for all three parts, plus limitations
README.md the evidence section

No secrets, no .env, no user data. proofs/out/ is gitignored by repo convention, so results are regenerated with the commands in README section 7.

mkthoma added 8 commits August 7, 2026 11:16
…live

Adds a 17-task cost-engineering workload, a second ladder config for a
two-arm crossover experiment, an adversarial task set, and the measured
evidence for all three assignment parts.

Ladder: the shipped frontier rung github/openai/gpt-4.1 was retired on
2026-07-30 -- the same day tiers.yaml recorded its measurement. Verified
dead across 17 model IDs, 7 vendors, plus the catalog endpoint. Repointed
to nvidia/nemotron-3-ultra-550b-a55b, priced from OpenRouter's published
rate for the identical model. Standard rung moved to gemini-3.5-flash-lite.

Measured findings:
- The ladder inverts. Projected 32.5x with frontier dearest; measured, the
  frontier rung costs LESS per call than economy, because economy writes
  7.3x more output tokens.
- The signature failure mode is observed (wide arm, B vs C): cost per call
  -34.5% while cost per resolved task rises +2.2%.
- Capability is not monotonic. The frontier rung resolved 1 of 5 moderate
  tasks where economy resolved 5 of 5, failing with arithmetic slips in
  5-24 output tokens. Cause: this policy's own reasoning:"off", set to
  avoid the invisible-token billing bug, which also removed the scratchpad.
- Truncation, not price, was the binding constraint. Raising the economy
  ceiling 512->1024 cut truncated attempts 18->4 and lifted resolution
  82.4%->94.1%, flipping the verdict.
- New attack: verbosity-induced escalation, 20.7x amplification per
  request, of which 6.45x is the defect rather than honest verbosity.

docs/LIMITATIONS.md records ten trace-derived limitations and nine
environmental ones, each marked OBSERVED or INFERRED -- including that p1
defaults to port 8112, silently falls back to a simulated transport, and
still reports ALL CHECKS PASSED with exit 0.
9/9 gates, live. Corroborates Part 2 on the shipped set: the frontier rung
resolved 11/12 where the cheapest resolved 12/12, and the budget-aware
cascade never left the economy rung, coming out cheapest per resolved task
at $0.00017956 -- 2.7x cheaper than always-frontier.

signature failure mode NOT OBSERVED here, which reproduces the session's own
published result for this task set on a rebuilt ladder.
README section 4 cited a Jaeger trace id, which is perishable evidence:
Jaeger all-in-one stores traces in memory and forgets them on restart, and
the /trace route mints a new id per call. Replaced with the raw
/api/traces/{id} response committed at docs/evidence/, verifiable without a
running backend. Swept for prompt/completion text before committing: none.

README section 5 spliced the verbosity attack together with p3's runaway
loop in a way that implied the attack itself had been refused. It had not --
every request is individually affordable, so all 12 were admitted. Added a
genuine before/after on the same attack: re-run under a $0.0006 per-task
ceiling, the two escalating tasks are refused at both standard and frontier,
spend falls 59.8% ($0.00627555 -> $0.00252000) and the controls are
untouched. The cost is stated rather than hidden: 12/12 resolved becomes
10/12. Also records that $0.0004 is too tight -- below the economy rung's
own projection, nothing runs at all.
The evidence section covered Parts 2 and 3 only; Part 1's results lived in
docs/PART1_EVIDENCE.md. Adds a Part 1 section to README.md directly: both
suite results with the two pre-existing failures explained, the six proofs
with p4's race characterised, the four evidence runs with prompt, tier,
model, decision, ledger and trace id, and the honest limitation the traces
exposed.
proofs/out/ is gitignored by repo convention, so the evidence sections cited
numbers with nothing to check them against. Adds the live results under
docs/evidence/ with an index: 11 proof outputs, the four Part 1 evidence run
captures, and one Jaeger trace as returned by its query API.

Simulated and duplicate runs omitted. Swept for keys, tokens and bearer
headers before committing; none present.
Puts the PR link at the top of the evidence section with a short description
of what it contains, plus a table routing each assignment part to its README
section, its detail document and the raw logs.
The evidence section was appended after 215 lines of existing documentation,
so the PR link sat at line 224 of 573 and nobody opening the file would see
it. Adds a callout box directly under the title with the PR link and jump
links to each part.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant