feat: implement compute workload budget policy, adversarial checks, and telemetry mapping - #9
Open
maniradhakrishnan-dev wants to merge 1 commit into
Conversation
…hecks, and unit tests
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This PR implements the budget and capability routing policies for Session 15 of EAGV3. We chose a custom domain workload for GPU Compute & AI Model Sizing consisting of 15 tasks of varying complexity, configured a custom 3-tier capability ladder, implemented hard budget controls, and verified them under adversarial runaway loop attacks.
What is implemented:
Part 1: Floor Reproduction
Part 2: Custom Policy & Workload Evaluation
config/tiers.yaml(Economy: GPT-OSS-120B, Standard: Gemini-3.1-Flash-Lite, Frontier: Gemini-3.1-Pro).proofs/run_compute_eval.py:$0.000480down to$0.000249).t09_flops_estimationwhere downgrading to economy saved$0.000566but lost correctness due to model math limitations.Part 3: Adversarial Budget Attack
proofs/attack_budget.pyto drive an infinite runaway planning loop under budget limits.$0.002ceiling) are neared, capping spend at$0.001927instead of$0.009241.task_failedinside the telemetry event trace.File Changes
config/tiers.yaml: Configured custom rungs and pricing.s15code/gateway.py: Added principal propagation and overrides passthrough filters.s15code/economics/controller.py: Enabled proper default fallback and principal mapping.proofs/tasks/compute_analysis.jsonl: The 15 custom model sizing tasks.proofs/run_compute_eval.py: Task runner and judge evaluation script.proofs/attack_budget.py: Runaway loop adversarial attack driver.proofs/capture_floor.py: Floor reproduction capture script.tests/test_cross_model_ladder.py: Adjusted model name assertion.README.md: Documented all Parts 1, 2, and 3 findings under the "Submission Evidence" section.