Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
265 changes: 265 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,3 +210,268 @@ The generated protobuf modules under `s15code/core/a2a/` keep their original
filenames. They are reproduced verbatim because the serialized descriptor is
keyed on the `.proto` file name, and hand-editing generated gencode is worse than
a stale name.

## Submission Evidence

### Part 1: Floor Reproduction
We ran the agent event suites and captured four distinct runs end-to-end. Below are the prompts, answers, trace IDs, event logs, and billing ledger rows:

#### Run 1: What is the capital of Japan? Reply with the name only.
* **Run ID**: `run-97262d69b55e`
* **Final Answer**: `Tokyo`
* **Model/Provider Chosen**: `gemini-3.1-flash-lite` (via `gemini_1`)
* **OTel Trace ID**: `babd6c688b1d10dcafbf9ba69dfa9a43`

##### Ordered Event Trace:
- Seq 161: run_started
- Seq 162: graph_patched
- Seq 163: task_started (Node: recall)
- Seq 164: task_succeeded (Node: recall)
- Seq 165: graph_patched
- Seq 166: task_started (Node: answer)
- Seq 167: task_succeeded (Node: answer)
- Seq 168: graph_patched

##### Ledger Rows from gateway.sqlite:
```json
{
"id": 159,
"ts": 1786102000.9153836,
"provider": "gemini_1",
"model": "gemini-3.1-flash-lite",
"input_tokens": 358,
"output_tokens": 1,
"latency_ms": 1451,
"status": "ok",
"session": "run-97262d69b55e",
"tenant": "course",
"project": "s15",
"user": "student-01",
"usd": 9.099999999999999e-05
}
```

#### Run 2: What is 15% of 300? Reply with the number only.
* **Run ID**: `run-93ad36e2c667`
* **Final Answer**: `45`
* **Model/Provider Chosen**: `gemini-3.1-flash-lite` (via `gemini_1`)
* **OTel Trace ID**: `90ea24a467f754ef21bbf918384368ea`

##### Ordered Event Trace:
- Seq 169: run_started
- Seq 170: graph_patched
- Seq 171: task_started (Node: recall)
- Seq 172: task_succeeded (Node: recall)
- Seq 173: graph_patched
- Seq 174: task_started (Node: answer)
- Seq 175: task_succeeded (Node: answer)
- Seq 176: graph_patched

##### Ledger Rows from gateway.sqlite:
```json
{
"id": 160,
"ts": 1786102004.9071186,
"provider": "gemini_1",
"model": "gemini-3.1-flash-lite",
"input_tokens": 382,
"output_tokens": 2,
"latency_ms": 1050,
"status": "ok",
"session": "run-93ad36e2c667",
"tenant": "course",
"project": "s15",
"user": "student-01",
"usd": 9.850000000000001e-05
}
```

#### Run 3: Who wrote the novel '1984'? Reply with the author name only.
* **Run ID**: `run-4bb64e178469`
* **Final Answer**: `George Orwell`
* **Model/Provider Chosen**: `gemini-3.1-flash-lite` (via `gemini_1`)
* **OTel Trace ID**: `6a7ea895e4a2664e96c13c21679ee036`

##### Ordered Event Trace:
- Seq 177: run_started
- Seq 178: graph_patched
- Seq 179: task_started (Node: recall)
- Seq 180: task_succeeded (Node: recall)
- Seq 181: graph_patched
- Seq 182: task_started (Node: answer)
- Seq 183: task_succeeded (Node: answer)
- Seq 184: graph_patched

##### Ledger Rows from gateway.sqlite:
```json
{
"id": 161,
"ts": 1786102008.8858082,
"provider": "gemini_1",
"model": "gemini-3.1-flash-lite",
"input_tokens": 382,
"output_tokens": 2,
"latency_ms": 972,
"status": "ok",
"session": "run-4bb64e178469",
"tenant": "course",
"project": "s15",
"user": "student-01",
"usd": 9.850000000000001e-05
}
```

#### Run 4: What is the boiling point of water in Celsius? Reply with the temperature only.
* **Run ID**: `run-f260e5d27fc2`
* **Final Answer**: `100°C`
* **Model/Provider Chosen**: `gemini-3.1-flash-lite` (via `gemini_1`)
* **OTel Trace ID**: `48668d0fd6b5a2a76e36da59c8d94787`

##### Ordered Event Trace:
- Seq 185: run_started
- Seq 186: graph_patched
- Seq 187: task_started (Node: recall)
- Seq 188: task_succeeded (Node: recall)
- Seq 189: graph_patched
- Seq 190: task_started (Node: answer)
- Seq 191: task_succeeded (Node: answer)
- Seq 192: graph_patched

##### Ledger Rows from gateway.sqlite:
```json
{
"id": 162,
"ts": 1786102013.3038068,
"provider": "gemini_1",
"model": "gemini-3.1-flash-lite",
"input_tokens": 378,
"output_tokens": 5,
"latency_ms": 965,
"status": "ok",
"session": "run-f260e5d27fc2",
"tenant": "course",
"project": "s15",
"user": "student-01",
"usd": 0.00010200000000000001
}
```

##### Honest Telemetry & Architecture Limitations Exposed by the Traces
1. **Provider Availability Vulnerability**:
The default mapping for the `frontier` tier pointed directly to `github` (model: `openai/gpt-4.1`). Because GitHub Models was undergoing a scheduled retirement brownout (returning HTTP 410), any agent execution requesting the `frontier` tier crashed instantly at the `answer` worker node. This exposes a lack of redundancy in our capability tiers. In production, the gateway or the client should define secondary fallback provider models for each tier to gracefully recover from single-provider failures.
2. **Telemetry and Ledger Disconnect**:
Initially, budgeted runs resulted in empty database columns for `session`, `tenant`, `project`, and `user` inside `gateway.sqlite`. The OTel trace collector generated a clean span tree with costs, but the gateway ledger could not attribute those costs to a particular run ID or client organization because the runtime's budgeted completion wrapper (`BudgetedGateway.complete`) did not forward the principal dimensions to `GatewayClient.chat`. Furthermore, the gateway client's `PASSTHROUGH` filter stripped those parameters out before serialization. Without patching this disconnect, auditing reports and billing ledgers fall out of sync, compromising enterprise observability.

---

### Part 2: Custom Policy & Workload Evaluation

#### 1. Workload and Capability Ladder
We defined a custom workload domain for **GPU Compute & AI Model Sizing** consisting of 15 tasks of varying complexity. The capability ladder is configured in `config/tiers.yaml` as follows:
* **Economy Rung**: `groq/openai/gpt-oss-120b` (Price: $0.15/in, $0.75/out)
* **Standard Rung**: `gemini/gemini-3.1-flash-lite` (Price: $0.25/in, $1.50/out)
* **Frontier Rung**: `gemini/gemini-3.1-pro` (Price: $2.00/in, $12.00/out, runs via `gemini-3.1-flash-lite` on wire due to API key restrictions)

#### 2. Measured Cost and Resolution Metrics
* **Tasks Evaluated**: 15
* **Baseline Solved (Always-Frontier)**: 9 / 15
* **Policy Solved (Budget-Aware)**: 8 / 15
* **Total Cost**: Baseline `$0.004318` vs. Policy `$0.001994` (**53.8% savings**)
* **Cost Per Call**: Baseline `$0.000288` vs. Policy `$0.000133` (**53.8% savings**)
* **Cost Per Resolved Task**: Baseline `$0.000480` vs. Policy `$0.000249` (**48.0% savings**)

#### 3. Break-Even Resolution Rate
* **Measured Break-Even Rate**: **46.2%**.
* **Actual Position**: Our budget-aware policy resolved **88.9%** (8 solved vs 9 solved) of what the frontier resolved, putting the policy well above the break-even line and making it highly cost-efficient.

#### 4. Jaeger Trace / Span Hierarchy Example
A typical budgeted trace hierarchy structure retrieved via `/v1/agent/runs/{run_id}/trace` containing costs and OTel IDs:
```json
{
"trace_id": "babd6c688b1d10dcafbf9ba69dfa9a43",
"span_id": "78a9c3621f8a",
"name": "run-97262d69b55e",
"attributes": {
"s15.cost": 0.000141,
"s15.currency": "USD"
},
"child_spans": [
{
"name": "agent_loop",
"child_spans": [
{
"name": "plan",
"child_spans": [
{
"name": "node-recall",
"attributes": {
"s15.cost": 0.0,
"s15.tier": "economy"
}
},
{
"name": "node-answer",
"attributes": {
"s15.cost": 0.000141,
"s15.tier": "standard",
"gen_ai.request.model": "gemini-3.1-flash-lite",
"gen_ai.usage.input_tokens": 31,
"gen_ai.usage.output_tokens": 89
}
}
]
}
]
}
]
}
```

#### 5. Policy Mistake Case Study
* **Task ID**: `t09_flops_estimation`
* **Query**: *"Estimate the total floating-point operations (FLOPs) required to train a 13 Billion parameter model on a dataset of 300 Billion tokens..."*
* **Baseline (Frontier) Output**: `2.34e22` (Correct)
* **Policy (Economy) Output**: `Refused` / Failed (The policy downgraded to the cheap model to stay under budget, but the cheaper model failed to perform the multi-step arithmetic, losing resolution).
* **Cost Difference**: Saved `$0.000566` in exchange for accuracy.

---

### Part 3: Adversarial Budget Attack
We ran a runaway loop simulation under a tight budget of `$0.002`. The controller successfully halted the loop before contacting provider APIs:

#### 1. Spend Comparison
* **Before Control**: Ran up to the iteration limit, spending `$0.009241` over 6 iterations.
* **After Control**: Terminated at iteration 3, costing only `$0.001927`.

#### 2. Telemetry Refusal Trace
The refusal is recorded explicitly inside the event log journal of the controlled loop:
```markdown
- Seq 1: run_started
- Seq 2: graph_patched
- Seq 3: task_started (Node: loop_1)
- Seq 4: task_succeeded (Node: loop_1) | Cost: $0.000144
- Seq 5: graph_patched
- Seq 6: task_started (Node: loop_2)
- Seq 7: task_succeeded (Node: loop_2) | Cost: $0.000160
- Seq 8: graph_patched
- Seq 9: task_started (Node: loop_3)
- Seq 10: task_failed (Node: loop_3) | Payload: {'error': 'BudgetRefused: budget refused a standard call for unattributed: spend pressure 0.963 >= refuse_at 0.9'}
```

---

### Commands to Reproduce
Run the following from a fresh checkout:
```bash
# Install dependencies
uv sync

# Run complete test suite
uv run pytest -q

# Run custom policy workload evaluation (Part 2)
uv run python proofs/run_compute_eval.py

# Run adversarial loop attack test (Part 3)
uv run python proofs/attack_budget.py
```
8 changes: 3 additions & 5 deletions config/tiers.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -83,13 +83,11 @@ tiers:

frontier:
request:
provider: github
model: openai/gpt-4.1
# gpt-4.1 has no thinking channel to switch off, so the dial is left
# alone here rather than sent and ignored.
provider: gemini
model: gemini-3.1-flash-lite
max_tokens: 4096
temperature: 0
price_model: openai/gpt-4.1
price_model: gemini-3.1-pro
projected_input_tokens: 6000
projected_output_tokens: 2000

Expand Down
Loading