Skip to content

feat: implement compute workload budget policy, adversarial checks, and telemetry mapping - #9

Open
maniradhakrishnan-dev wants to merge 1 commit into
theschoolofai:mainfrom
maniradhakrishnan-dev:feature/compute-budget-policy
Open

maniradhakrishnan-dev wants to merge 1 commit into
theschoolofai:mainfrom
maniradhakrishnan-dev:feature/compute-budget-policy

Conversation

@maniradhakrishnan-dev

Copy link
Copy Markdown

Overview

This PR implements the budget and capability routing policies for Session 15 of EAGV3. We chose a custom domain workload for GPU Compute & AI Model Sizing consisting of 15 tasks of varying complexity, configured a custom 3-tier capability ladder, implemented hard budget controls, and verified them under adversarial runaway loop attacks.


What is implemented:

Part 1: Floor Reproduction

  • Ran the agent event suites and captured event logs, SQLite database records, and OpenTelemetry trace IDs for 4 distinct runs.
  • Documented key limitations regarding provider availability vulnerabilities and telemetry-ledger disconnects.

Part 2: Custom Policy & Workload Evaluation

  • Configured a custom capability ladder in config/tiers.yaml (Economy: GPT-OSS-120B, Standard: Gemini-3.1-Flash-Lite, Frontier: Gemini-3.1-Pro).
  • Built a budget controller that propagates tenant/project/user principal tags to the gateway and performs automated tier downgrades.
  • Evaluated performance using a Gemini judge in proofs/run_compute_eval.py:
    • Cost Savings: Reduced cost per resolved task by 48% (from $0.000480 down to $0.000249).
    • Break-Even Rate: Evaluated the ladder's break-even resolution rate at 46.2%.
    • Mistake Analysis: Analyzed task t09_flops_estimation where downgrading to economy saved $0.000566 but lost correctness due to model math limitations.

Part 3: Adversarial Budget Attack

  • Created proofs/attack_budget.py to drive an infinite runaway planning loop under budget limits.
  • Proven that the controller hard-refuses provider calls once budget limits (e.g. $0.002 ceiling) are neared, capping spend at $0.001927 instead of $0.009241.
  • Refusals are fully logged as task_failed inside the telemetry event trace.

File Changes

  • config/tiers.yaml: Configured custom rungs and pricing.
  • s15code/gateway.py: Added principal propagation and overrides passthrough filters.
  • s15code/economics/controller.py: Enabled proper default fallback and principal mapping.
  • proofs/tasks/compute_analysis.jsonl: The 15 custom model sizing tasks.
  • proofs/run_compute_eval.py: Task runner and judge evaluation script.
  • proofs/attack_budget.py: Runaway loop adversarial attack driver.
  • proofs/capture_floor.py: Floor reproduction capture script.
  • tests/test_cross_model_ladder.py: Adjusted model name assertion.
  • README.md: Documented all Parts 1, 2, and 3 findings under the "Submission Evidence" section.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant