Conversation
…live Adds a 17-task cost-engineering workload, a second ladder config for a two-arm crossover experiment, an adversarial task set, and the measured evidence for all three assignment parts. Ladder: the shipped frontier rung github/openai/gpt-4.1 was retired on 2026-07-30 -- the same day tiers.yaml recorded its measurement. Verified dead across 17 model IDs, 7 vendors, plus the catalog endpoint. Repointed to nvidia/nemotron-3-ultra-550b-a55b, priced from OpenRouter's published rate for the identical model. Standard rung moved to gemini-3.5-flash-lite. Measured findings: - The ladder inverts. Projected 32.5x with frontier dearest; measured, the frontier rung costs LESS per call than economy, because economy writes 7.3x more output tokens. - The signature failure mode is observed (wide arm, B vs C): cost per call -34.5% while cost per resolved task rises +2.2%. - Capability is not monotonic. The frontier rung resolved 1 of 5 moderate tasks where economy resolved 5 of 5, failing with arithmetic slips in 5-24 output tokens. Cause: this policy's own reasoning:"off", set to avoid the invisible-token billing bug, which also removed the scratchpad. - Truncation, not price, was the binding constraint. Raising the economy ceiling 512->1024 cut truncated attempts 18->4 and lifted resolution 82.4%->94.1%, flipping the verdict. - New attack: verbosity-induced escalation, 20.7x amplification per request, of which 6.45x is the defect rather than honest verbosity. docs/LIMITATIONS.md records ten trace-derived limitations and nine environmental ones, each marked OBSERVED or INFERRED -- including that p1 defaults to port 8112, silently falls back to a simulated transport, and still reports ALL CHECKS PASSED with exit 0.
9/9 gates, live. Corroborates Part 2 on the shipped set: the frontier rung resolved 11/12 where the cheapest resolved 12/12, and the budget-aware cascade never left the economy rung, coming out cheapest per resolved task at $0.00017956 -- 2.7x cheaper than always-frontier. signature failure mode NOT OBSERVED here, which reproduces the session's own published result for this task set on a rebuilt ladder.
README section 4 cited a Jaeger trace id, which is perishable evidence:
Jaeger all-in-one stores traces in memory and forgets them on restart, and
the /trace route mints a new id per call. Replaced with the raw
/api/traces/{id} response committed at docs/evidence/, verifiable without a
running backend. Swept for prompt/completion text before committing: none.
README section 5 spliced the verbosity attack together with p3's runaway
loop in a way that implied the attack itself had been refused. It had not --
every request is individually affordable, so all 12 were admitted. Added a
genuine before/after on the same attack: re-run under a $0.0006 per-task
ceiling, the two escalating tasks are refused at both standard and frontier,
spend falls 59.8% ($0.00627555 -> $0.00252000) and the controls are
untouched. The cost is stated rather than hidden: 12/12 resolved becomes
10/12. Also records that $0.0004 is too tight -- below the economy rung's
own projection, nothing runs at all.
The evidence section covered Parts 2 and 3 only; Part 1's results lived in docs/PART1_EVIDENCE.md. Adds a Part 1 section to README.md directly: both suite results with the two pre-existing failures explained, the six proofs with p4's race characterised, the four evidence runs with prompt, tier, model, decision, ledger and trace id, and the honest limitation the traces exposed.
proofs/out/ is gitignored by repo convention, so the evidence sections cited numbers with nothing to check them against. Adds the live results under docs/evidence/ with an index: 11 proof outputs, the four Part 1 evidence run captures, and one Jaeger trace as returned by its query API. Simulated and duplicate runs omitted. Swept for keys, tokens and bearer headers before committing; none present.
Puts the PR link at the top of the evidence section with a short description of what it contains, plus a table routing each assignment part to its README section, its detail document and the raw logs.
The evidence section was appended after 215 lines of existing documentation, so the PR link sat at line 224 of 573 and nobody opening the file would see it. Adds a callout box directly under the title with the PR link and jump links to each part.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A budget-aware routing policy for a LLM / cloud cost engineering workload, with the measurement, the break-even analysis, and the places it goes wrong. All numbers measured live 2026-08-07 against
glc_v4.Jump links are in a box at the top of
README.md. The evidence sections arePart 1, then sections 1-7 for Parts 2 and 3; full working in
docs/, and theraw live results are committed under
docs/evidence/.The ladder had to be rebuilt first
config/tiers.yamlrecords its three rungs as measured on 2026-07-30. GitHub Models - which served the frontier rung,openai/gpt-4.1- was fully retired on 2026-07-30, the same day.Verified dead on 2026-08-06: 17 model IDs across 7 vendors all return
HTTP 410 github_models_retirement_brownout, as doesGET /catalog/modelsand a request naming no model at all. A deliberately invalid ID returns 410 rather than 404, proving the endpoint short-circuits before model resolution.openai/gpt-oss-120bgemini-3.5-flash-litenvidia/nemotron-3-ultra-550b-a55bFrontier priced from OpenRouter's published rate for the identical model, since NVIDIA's free tier bills $0 and a free top rung makes "refuse at exhaustion" inexpressible in dollars.
Measured, against an always-frontier baseline
Two arms, identical but for the token ceiling - same tasks, models, prices, judge.
Wide ladder (economy ceiling 512):
The signature failure mode was observed (B vs C): cost per call -34.5% while cost per resolved task rose +2.2%. Change the denominator, reverse the ranking.
The measured ladder is inverted. Projected 32.5x with frontier dearest; measured, frontier costs less per call than economy - the economy model writes 7.3x more output tokens (407.6 vs 55.6 mean), which more than cancels frontier's 4.8x output rate.
Break-even
r* = k / (C/c - 1 + k), k = 3. In the wide arm the cheap rung sat 3.3 points below break-even (0.824 against 0.8562), which is why it lost.What moved it was not price - it was truncation. Raising the economy ceiling 512->1024 cut truncated attempts from 18 to 4 and lifted resolution 82.4% -> 94.1%, carrying it well above break-even. Strategy A is the control at 10/17 in both arms, since it writes 5-24 output tokens and never nears a ceiling.
My original hypothesis - that a narrower price spread would make the trap appear - was wrong, and the confound is stated plainly in the docs: equalising
max_tokenschanged the ceiling and the projected spread together, because the projection is a function of the ceiling.Where the policy chose wrongly
c06_cache_discount:Same model, same prompt, temperature 0 - one run wrote 251 tokens and finished, the other ran to exactly 512 and was cut off. The cascade cannot distinguish "the model was wrong" from "the model ran out of room."
Reported beside
c11, where the same rule climbed all three rungs at 2.8x and reached an answer B never did.Capability is not monotonic. The frontier rung resolved 1 of 5 moderate tasks where economy resolved 5 of 5, failing with arithmetic slips in 5-24 output tokens (
0.018360for0.018346;$3.30for$4.30). The cause is this policy's ownreasoning: "off"- set to avoid the invisible-token billing bug, it also removed the model's scratchpad. The dial that stops you paying for invisible tokens is the same dial that costs you correctness on multi-step work.Every frontier miss scored exactly 0.6667 - 4 on all four generic criteria, 0 on
meets_expectation. Precisely the case the 0.5 per-criterion floor exists to catch.The adversarial test
proofs/tasks/verbosity_attack.jsonl- six requests whose correct answers are long, paired with six controls from the same domain that are equally easy but short. Every request is individually cheap and passes any spend ceiling.20.7x amplification per request, all 12 resolved - it inflates cost while looking like a completely successful workload. The escalation multiplier of 6.45x independently reproduces the 7.0x measured on
c06. One request missed the cap by six tokens; asking for "integers 1 to 400" makes it deterministic.Before/after the control (
p3, 60 rounds, $0.0005): 1 admitted, 59 refused, 59 nodes failed withBudgetRefused.Limitations found while reproducing the floor
docs/LIMITATIONS.md- ten trace-derived, nine environmental, each marked OBSERVED or INFERRED. The ones that may matter to this repo:p1silently simulates and reports success. It defaults to port 8112 while every other proof andGLC_PORTuse 8111, and no proof loads.env, soGLC_BASE_URLis never read on that path. A live-intended run against a healthy gateway produced a full cost table, a break-even calculation and asignature failure mode OBSERVEDfinding - all simulated, exit 0,ok: true, every gate green.p4's backend gates race ingestion - 8 of 10 spans on one run, 7 of 10 on the next; re-querying seconds later returns all 10. No wait or retry between export and assertion.p4's query URL is unreachable on Windows/WSL - Jaeger binds[::]and WSL will not relay that to the IPv4 literal.S15_TRACE_QUERY_URL=http://localhost:16686works./v1/agent/runs/{id}/tracere-exports and mints a new one per call; Jaeger held 11 traces for 5 runs, one run in 5 copies.provider_calls=0, cost=0on a run that made a real billable call.GEMINI_MODELdefaults to the retiredgemini-2.5-flash, which fails every unbudgeted run.jaeger_local.shrejects Windows despite Jaeger shipping a Windows binary, and ships CRLF that breaks it under WSL.file://assumption).What's in the diff
proofs/tasks/llm_cost_engineering.jsonlproofs/tasks/verbosity_attack.jsonlconfig/tiers.yaml,config/pricing.yamlconfig_narrow/docs/*.mdREADME.mdNo secrets, no
.env, no user data.proofs/out/is gitignored by repo convention, so results are regenerated with the commands in README section 7.