Skip to content

feat: session 15 assignment - code review policy, proofs & evidence - #13

Open
manishmcsa01-cmd wants to merge 1 commit into
theschoolofai:mainfrom
manishmcsa01-cmd:submission
Open

manishmcsa01-cmd wants to merge 1 commit into
theschoolofai:mainfrom
manishmcsa01-cmd:submission

Conversation

@manishmcsa01-cmd

@manishmcsa01-cmd manishmcsa01-cmd commented Aug 8, 2026 •

Copy link
Copy Markdown

Session 15 Assignment — Model Routing, Agent Economics & Observability

Author: Manish Kapoor (manishmcsa01-cmd)
Workload Domain: Automated Code Review & Security Audit (proofs/tasks/code_review.jsonl)


Key Evidence & Findings Summary

1. Floor Reproduction & Unit Tests

  • glc_v4 test suite: 436 passed, 20 skipped
  • S15Code test suite: 277 passed, 0 failed
  • All core proof scripts executed cleanly against live gateway (GLC_BASE_URL=http://127.0.0.1:8111):
    • p2_budget_holds: PASS (5/5 checks)
    • p3_denial_of_wallet: PASS (6/6 checks)
    • p4_trace_export: PASS (11/11 checks) — Trace ID: 4db4392e22b2440f468bdb4813b7c9b0
    • p7_cross_model_ladder: PASS (10/10 checks) — 3 distinct providers/models verified.

2. Custom Workload & Economics

  • Capability Ladder:
    • Economy: groq / openai/gpt-oss-120b ($0.15 / $0.75 per Mtok)
    • Standard: gemini / gemini-3.1-flash-lite ($0.25 / $1.50 per Mtok)
    • Frontier: groq / llama-3.3-70b-versatile ($1.50 / $4.50 per Mtok)
  • Disjoint Judge Panel: gemini-2.5-flash and meta/llama-3.1-8b-instruct.
  • Cost-per-Call Fallacy: Strategy B (Always Economy) lowered cost per call by 63.1%, but resulted in a 10.7% HIGHER cost per resolved task ($0.00008432 vs $0.00007616) due to retries on complex code review tasks.
  • Winner: Strategy C (Budget-Aware Cascade) achieved an 87.5% resolution rate while cutting cost per resolved task by 41.7% compared to Strategy A.
  • Break-Even Resolution Rate: Derived at 36.9% from measured price spread.

3. Adversarial Budget Attack (proofs/adversarial_budget.py)

  • ALL 4 SCENARIOS PASSED ✓:
    • Scenario A (Runaway Loop): BudgetRefused triggered and spend capped.
    • Scenario B (Unaffordable Tier): Micro-budget ($0.000001) refused prior to provider call.
    • Scenario C (Cascade Climb): Escalation capped by budget ceiling.
    • Scenario D (Metering Verification): Transport calls match ledger charges rendering exact agreement.

Full benchmark tables, proof JSON outputs, case studies, and reproduction commands are documented in README.md.

@manishmcsa01-cmd

Copy link
Copy Markdown
Author

Session 15 Assignment — Model Routing, Agent Economics & Observability

Author: Manish Kapoor (manishmcsa01-cmd)
Workload Domain: Automated Code Review & Security Audit (proofs/tasks/code_review.jsonl)


Key Evidence & Findings Summary

1. Floor Reproduction & Unit Tests

  • glc_v4 test suite: 436 passed, 20 skipped
  • S15Code test suite: 277 passed, 0 failed
  • All core proof scripts executed cleanly against live gateway (GLC_BASE_URL=http://127.0.0.1:8111):
    • p2_budget_holds: PASS (5/5 checks)
    • p3_denial_of_wallet: PASS (6/6 checks)
    • p4_trace_export: PASS (11/11 checks) — Trace ID: 4db4392e22b2440f468bdb4813b7c9b0
    • p7_cross_model_ladder: PASS (10/10 checks) — 3 distinct providers/models verified.

2. Custom Workload & Economics

  • Capability Ladder:
    • Economy: groq / openai/gpt-oss-120b ($0.15 / $0.75 per Mtok)
    • Standard: gemini / gemini-3.1-flash-lite ($0.25 / $1.50 per Mtok)
    • Frontier: groq / llama-3.3-70b-versatile ($1.50 / $4.50 per Mtok)
  • Disjoint Judge Panel: gemini-2.5-flash and meta/llama-3.1-8b-instruct.
  • Cost-per-Call Fallacy: Strategy B (Always Economy) lowered cost per call by 63.1%, but resulted in a 10.7% HIGHER cost per resolved task ($0.00008432 vs $0.00007616) due to retries on complex code review tasks.
  • Winner: Strategy C (Budget-Aware Cascade) achieved an 87.5% resolution rate while cutting cost per resolved task by 41.7% compared to Strategy A.
  • Break-Even Resolution Rate: Derived at 36.9% from measured price spread.

3. Adversarial Budget Attack (proofs/adversarial_budget.py)

  • ALL 4 SCENARIOS PASSED ✓:
    • Scenario A (Runaway Loop): BudgetRefused triggered and spend capped.
    • Scenario B (Unaffordable Tier): Micro-budget ($0.000001) refused prior to provider call.
    • Scenario C (Cascade Climb): Escalation capped by budget ceiling.
    • Scenario D (Metering Verification): Transport calls match ledger charges exactly.

Full benchmark tables, proof JSON outputs, case studies, and reproduction commands are documented in README.md.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant