Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
305 changes: 305 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,3 +210,308 @@ The generated protobuf modules under `s15code/core/a2a/` keep their original
filenames. They are reproduced verbatim because the serialized descriptor is
keyed on the `.proto` file name, and hand-editing generated gencode is worse than
a stale name.

---

<a id="part-1"></a>
## Assignment evidence — Part 1: reproduce the floor

All evidence below was produced on 1 September 2026 by the live gateway and
local models. It is committed under [`docs/evidence/`](docs/evidence/) as JSON,
not copied from terminal output.

### Important pricing disclosure

The models run locally through Ollama, so no provider billed dollars. Dollar
figures use the explicit **declared rate card** in
[`config/pricing.yaml`](config/pricing.yaml): $0.10/$0.40 per million tokens for
the 1B rung, $0.30/$1.20 for the 3B rung, and $1.00/$4.00 for the 7B rung.
Token counts, latency, routing, resolution verdicts, and refusals are measured;
only conversion from tokens to dollars is synthetic. Reporting these values as
real provider charges would be false.

The gateway extension that registers the second local slot is in the public
[`tanmays369/glc_v4`](https://github.com/tanmays369/glc_v4) fork. `ollama` and
`ollama_edge` are separate daemons, queues, rate states, and circuit breakers,
but they run on one machine and are not independent vendors.

### Test and proof floor

- glc_v4: **456 passed, 1 skipped**
- S15Code: **277 passed**
- `p1`: all checks passed over the shipped 12-task file; 59 answering calls
equalled 59 ledger charges, with zero transport failures and zero unusable
judge samples ([result](docs/evidence/p1_cost_per_task_floor.json)).
- `p2`: budget held, a tight allowance downgraded frontier to standard, and an
impossible allowance refused before transport
([result](docs/evidence/p2_budget_holds.json)).
- `p3`: a loop kept asking for 80 rounds; 23 calls were admitted, 57 refused,
and every refusal was a visible `BudgetRefused` graph failure
([result](docs/evidence/p3_denial_of_wallet.json)).
- `p4`: the exported provider-call cost was **$0.00106400**, exactly equal to
the ledger (delta zero); the trace was fetched back from Jaeger and contained
no prompt or completion text ([result](docs/evidence/p4_trace_export.json)).
- `p7`: every rung served its configured model; both downgrade steps changed
the model. Projected spread was 39.3x and measured spread was 3.82x
([result](docs/evidence/p7_cross_model_ladder.json)).

### Four captured runs

The common ordered event shape was:

`run_started → graph_patched → task_started(recall) → task_succeeded(recall) → graph_patched → task_started(answer) → task_succeeded|task_failed(answer) → graph_patched`

Every run patches in a `recall` node at economy before the `answer` node it was
asked for, so the tier the graph requests and the tier the ledger charges are
both recorded below.

**Run 1 — CrashLoopBackOff.** Trace `cc3fc87ce212ddadf43d5c0698b43937`.

> A pod is in CrashLoopBackOff. Give the first three kubectl commands you would
> run, in order, using placeholders my-ns and my-pod.

| seq | node | requested tier | served | in | out | projected | charged |
|---:|---|---|---|---:|---:|---:|---:|
| 1 | `answer` | frontier | `ollama_edge/qwen2.5:7b` | 248 | 208 | $0.008681 | **$0.001080** |

Run spend $0.001080 over 1 call, 0 refusals. `recall` ran at economy and billed
nothing. The final answer returned `kubectl get pods`, `kubectl describe pod`
and `kubectl logs`, in that order.
[Run, journal, spans](docs/evidence/run1_crashloop.json)
· [Jaeger response](docs/evidence/jaeger_cc3fc87ce212ddadf43d5c0698b43937.json)

**Run 2 — ImagePullBackOff.** Trace `7654bdf03b8fbfac757cd2026a726be9`.

> A Pod is stuck in ImagePullBackOff. List the first three places you inspect,
> in order.

| seq | node | requested tier | served | in | out | projected | charged |
|---:|---|---|---|---:|---:|---:|---:|
| 1 | `answer` | frontier | `ollama_edge/qwen2.5:7b` | 507 | 138 | $0.009067 | **$0.001059** |

Run spend $0.001059 over 1 call, 0 refusals. The answer named the same three
`kubectl` steps rather than the registry, image tag and pull-secret the question
was reaching for — a correct-shaped answer to a slightly different question.
[Run, journal, spans](docs/evidence/run2_imagepull.json)
· [Jaeger response](docs/evidence/jaeger_7654bdf03b8fbfac757cd2026a726be9.json)

**Run 3 — HTTP status for a budget refusal.** Trace
`dceacc434b42228f416605cf4fc59048`.

> An LLM gateway refuses a call because the remaining spend budget cannot cover
> the projected cost. What HTTP status code should it return? Reply with the
> number only.

| seq | node | requested tier | served | in | out | projected | charged |
|---:|---|---|---|---:|---:|---:|---:|
| 1 | `answer` | frontier | `ollama_edge/qwen2.5:7b` | 719 | 4 | $0.009371 | **$0.000735** |

Run spend $0.000735 over 1 call, 0 refusals. Final answer: `402`. Note the
projection was 12.7x the charge here — the estimator sizes the output envelope
it must reserve, and a four-token reply leaves most of that reservation unspent.
[Run, journal, spans](docs/evidence/run3_http402.json)
· [Jaeger response](docs/evidence/jaeger_dceacc434b42228f416605cf4fc59048.json)

**Run 4 — unaffordable call under a $0.00003 ceiling.** Trace
`f86dcc101182b9ebc60153f048965918`.

> Compare two approaches to bounding agent spend in three sentences.

No ledger rows: the run made zero provider calls and spent $0.00. The `answer`
node failed with `BudgetRefused` and the refusal record reads:

```json
{"node_id": "answer", "action": "refuse", "requested_tier": "frontier",
"tier": null, "model": null, "projected_cost": 0.0003273,
"spent": 0.0, "remaining": 0.00003, "allowance": 0.000024, "ladder_steps": 0,
"reason": "cheapest tier economy projects 0.000327, run holds 0.000030 (headroom 0.000001)"}
```

`ladder_steps: 0` is the load-bearing detail: the controller did not try
frontier, fail, and walk down. It priced the *cheapest* rung first, found even
that unaffordable, and refused before any transport. The final answer is the
empty string, because no model was ever asked.
[Run, journal, spans](docs/evidence/run4_tight_budget.json)
· [Jaeger response](docs/evidence/jaeger_f86dcc101182b9ebc60153f048965918.json)

The complete compact capture is
[`four_runs_summary.json`](docs/evidence/four_runs_summary.json), and the
gateway ledger extract is
[`gateway_ledger_recent.json`](docs/evidence/gateway_ledger_recent.json).

![Jaeger span waterfall](docs/evidence/jaeger_trace.png)

The screenshot is the actual Jaeger UI opened by trace ID; the compact
[`SVG waterfall`](docs/evidence/jaeger_waterfall.svg) and raw
[`query response`](docs/evidence/jaeger_cc3fc87ce212ddadf43d5c0698b43937.json)
are committed so the image can be checked against backend data.

### Honest limitation exposed by the traces

The refused run has nine agent spans but no provider-call span, because no
provider call existed. The refusal is visible in the graph journal and budget
`refusal_log`, not as a billed model span. An operator looking only for provider
spans would therefore miss the system's safest behavior. A production dashboard
must ingest refusal events as a separate control signal.

<a id="part-2"></a>
## Assignment evidence — Part 2: SRE routing policy

The workload is 18 self-contained platform/SRE triage tasks in
[`proofs/tasks/sre_triage.jsonl`](proofs/tasks/sre_triage.jsonl): Kubernetes API
versions, Pod diagnosis, Jenkins investigation order, Jira JQL, IDP stack
nesting, budget arithmetic, routing pools, caching, and billed-empty-output
diagnosis. Difficulty mix: 4 trivial, 6 moderate, 8 hard.

The ladder and role policy are data in [`config/tiers.yaml`](config/tiers.yaml):

- economy — `ollama/llama3.2:1b`, max 512 tokens
- standard — `ollama/qwen2.5:3b`, max 1,024 tokens
- frontier — `ollama_edge/qwen2.5:7b`, max 2,048 tokens

Bulk retrieval and formatting request economy, operational synthesis requests
standard, and the final answer node requests frontier. The hard controller in
[`config/budgets.yaml`](config/budgets.yaml) downgrades at 50% pressure, refuses
at 90%, reserves 20% for the terminal node, and caps a run at 60 calls.

The judge uses the generic five-criterion rubric from
[`config/evals.yaml`](config/evals.yaml). Four criteria have weight 1 and the
task-specific success criterion has weight 2. Resolution requires a weighted
score of at least 0.75 and no individual criterion below 0.5. The floor matters:
a fluent, specific, complete, internally consistent answer to the wrong question
scores 4/6 = 0.667 on generic criteria and cannot pass. `phi4-mini` is disjoint
from all answering models, but it is a single local judge, so inter-judge
agreement cannot be measured and its verdicts have that explicit limitation.

<!-- SRE_RESULTS_START -->
The live measurement completed in 1,019.2 seconds. All 90 answering calls that
returned had matching ledger charges; 8 provider failures reported no usage and
were kept separate rather than assigned invented costs. Full output:
[`p1_cost_per_task_sre.json`](docs/evidence/p1_cost_per_task_sre.json) and
[`p1_sre_run.log`](docs/evidence/p1_sre_run.log).

| Strategy | Spend | Calls | Cost/call | Resolved | Cost/resolved task |
|---|---:|---:|---:|---:|---:|
| A: always frontier | $0.00882300 | 18 | $0.00049017 | 9/18 | $0.00098033 |
| B: always economy, two retries | $0.00290220 | 42 | $0.00006910 | 6/18 | $0.00048370 |
| C: budget-aware cascade | $0.00349260 | 30 | $0.00011642 | 10/18 | $0.00034926 |

Against always-frontier, the policy reduced cost per call by **76.25%**, reduced
cost per resolved task by **64.37%**, and resolved one additional task in this
single run. Those percentages are consequences of the declared rate card, not
provider invoices.

### Break-even

The cheap strategy used at most `k=3` attempts. From its measured cost/call and
the frontier's measured cost/resolved task:

`r* = k / (C/c - 1 + k) = 18.5%`

Economy resolved 33.3%, leaving 14.8 percentage points of measured headroom
above the frontier break-even rate. Against the cascade, however, its break-even
was 42.5%; economy sat below it. This is the session's signature trap: B was
40.6% cheaper per call than C but **38.5% more expensive per resolved task**.

### Where the policy chose wrongly

On `s11_eks_api_table`, the frontier baseline answered all five Kubernetes API
versions correctly and resolved for **$0.00029000**. The policy opened on
economy, which incorrectly put Deployment, Job, and Ingress on `v1`, then paid
standard for invented software-like versions such as `2.5.0`. Those two wrong
attempts cost **$0.00013320**. Its frontier escalation then hit a real 120-second
`ReadTimeout`, so the task remained unresolved.

This is both a routing failure and an operational limitation: the cheap-first
policy bought two unusable answers before asking the only rung that had
demonstrated success, while the single local frontier slot had no failover when
it timed out. It would be dishonest to attribute the final failure solely to
model quality; the evidence records both causes.

### Judging cost

Evaluation used 46 judge calls costing **$0.01648380** on the same declared card:
4.72x the policy workload it judged and 1.87x the always-frontier workload.
There were zero unparseable verdicts and 44 memoized verdict reuses. The
meta-cost is larger than the work, so an online router should not call this
judge on every production request.
<!-- SRE_RESULTS_END -->

<a id="part-3"></a>
## Assignment evidence — Part 3: attack the budget

[`proofs/p8_attack_policy.py`](proofs/p8_attack_policy.py) attacks the policy in
four ways and exits non-zero unless every control holds
([machine-readable result](docs/evidence/p8_attack_policy.json)).

- **Runaway loop:** under a $0.002 ceiling, spend flattened at **$0.001764**;
the loop reached 80 rounds and produced **38 visible refusals**. With a loose
ceiling, all eight sampled rounds kept spending. Extrapolating that measured
rate to 10,000 rounds gives about **$1.61** on the declared card.
- **Unaffordable frontier:** projected $0.00021060 against $0.00000029
remaining; shortfall $0.00021031. Calls: 0, spend: 0, refusals: 1.
- **Forced full cascade:** the same task climbed
`llama3.2:1b → qwen2.5:3b → qwen2.5:7b` and spent **$0.00037440**. A wrong
confidence signal buys every rung it passes.
- **Failure after token consumption:** the successful call was charged; two
subsequent errors with no reported usage were counted as transport failures
and not invented as charges. A fresh exhausted request was refused and its
envelope remained in `refusal_log`.

This demonstrates the distinction the policy relies on: billed calls become
ledger charges and provider spans; pre-call refusals become explicit graph
failures and refusal records.

## Reproduce from a fresh checkout

Prerequisites: Python 3.11+, `uv`, Ollama, and about 11 GB for the four local
models. No cloud API key is needed.

```bash
git clone https://github.com/tanmays369/glc_v4.git
git clone https://github.com/tanmays369/S15Code.git

ollama pull llama3.2:1b
ollama pull qwen2.5:3b
ollama pull qwen2.5:7b
ollama pull phi4-mini

# Terminal 1: stock Ollama daemon
ollama serve

# Terminal 2: independent local provider slot
OLLAMA_HOST=127.0.0.1:11435 ollama serve

# Terminal 3: Jaeger without Docker (use docker-compose.observability.yml if preferred)
cd glc_v4
scripts/jaeger_local.sh --start

# Copy the assignment gateway policy and start on 8121.
mkdir -p runtime/glc
cp examples/s15/pricing.yaml runtime/glc/pricing.yaml
cp examples/s15/routing.yaml runtime/glc/routing.yaml
export GLC_CONFIG_DIR="$PWD/runtime/glc"
export OLLAMA_MODEL=qwen2.5:3b
export OLLAMA_URL=http://127.0.0.1:11434
export OLLAMA_EDGE_MODEL=qwen2.5:7b
export OLLAMA_EDGE_URL=http://127.0.0.1:11435
export LLM_ORDER=ollama,ollama_edge
export OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:4318
uv sync
uv run glc serve --port 8121

# Terminal 4: runtime and proofs
cd ../S15Code
export GLC_BASE_URL=http://127.0.0.1:8121
export S15_OTEL_EXPORTER_ENDPOINT=http://127.0.0.1:4318/v1/traces
uv sync
uv run pytest -q
uv run python proofs/p2_budget_holds.py --task "Diagnose CrashLoopBackOff" --budget 0.02
uv run python proofs/p3_denial_of_wallet.py --task "Diagnose CrashLoopBackOff" --budget 0.001 --loop-limit 80
uv run python proofs/p4_trace_export.py --task "Diagnose CrashLoopBackOff" --budget 0.02 \
--otel-endpoint http://127.0.0.1:4318/v1/traces
uv run python proofs/p7_cross_model_ladder.py --task "Diagnose CrashLoopBackOff"
uv run python proofs/p8_attack_policy.py --task "Diagnose CrashLoopBackOff" --budget 0.002
uv run python proofs/p1_cost_per_task.py --tasks proofs/tasks/sre_triage.jsonl \
--principal course/s15/reviewer --label sre
```
29 changes: 23 additions & 6 deletions config/budgets.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,18 @@
# agent — whatever string the caller attributes spend to). The controller is
# deterministic code, not a prompt: token elasticity means a model asked nicely
# to stay under a ceiling will sail straight through it.
#
# Tuned for the platform/SRE triage workload in proofs/tasks/sre_triage.jsonl.
# Worst-case projected cost per call on the ladder in tiers.yaml:
#
# economy 900 in x 0.10 + 512 out x 0.40 = $0.00029480
# standard 1400 in x 0.30 + 1024 out x 1.20 = $0.00165120
# frontier 2200 in x 1.00 + 2048 out x 4.00 = $0.01039200
#
# a 35x projected spread, which is what makes "one rung down" worth doing.

# Used when a caller asks for a budgeted run without naming an amount.
# Used when a caller asks for a budgeted run without naming an amount. Four
# frontier calls, or roughly thirty economy ones.
default_budget: 0.05

# Fraction of the run's allowance held back from frontier allocation so the
Expand All @@ -26,13 +36,20 @@ headroom_fraction: 0.02

# Admission prices the WORST case, not a guess: output is bounded by the tier's
# max_tokens and input by the prompt actually being sent. These two numbers turn
# the prompt into a token count — roughly four characters per token, then a safety
# factor because a provider's tokeniser is not ours.
chars_per_token: 4
input_estimate_safety: 1.25
# the prompt into a token count.
#
# 4 chars/token is the usual English prose figure and it is WRONG for this task
# class. These prompts carry YAML specs, kubectl invocations, JQL and stack
# traces, all of which tokenise denser than prose — punctuation, colons, dashes
# and CamelCase each cost a token. Estimating at 3.2 chars/token makes the
# estimate pessimistic in the direction that matters: over-counting input makes
# admission stricter, and a budget controller that under-counts is not one.
chars_per_token: 3.2
input_estimate_safety: 1.3

# Hard ceiling on admitted calls per run. Denial-of-wallet defence that does not
# depend on price estimates being right.
# depend on price estimates being right — the one control that still holds when
# every price on the card is wrong.
max_calls_per_run: 60

# Hard ceiling on admitted calls attributable to one graph node, so a single
Expand Down
Loading