S15 Parts 2 & 3: a ticket-triage policy, measured against always-frontier and attacked - Dheeraj Hegde - #10
Open
Dheeraj-Hegde wants to merge 3 commits into
Open
Dheeraj-Hegde wants to merge 3 commits into
Dheeraj-Hegde wants to merge 3 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Builds a new task class, its ladder and its budget policy on top of Part 1's
runtime, measures the policy against an always-frontier baseline, and then
attacks it. Nothing in
s15code/is edited — the whole submission is one newtask file, one new config directory, two proof scripts and evidence under
proofs/out/. The main README.md has the full sections; twostandalone review docs are added for convenience:
Headline numbers
Part 2 — 18 IT-ticket-triage tasks
(proofs/tasks/tickets.jsonl) through the
Part 2 ladder and policy (config/part2/):
resolution.
2 attempts: 12.8 %. B sits at 33 % — above break-even, so B really is
cheaper per resolved on the tasks it does resolve, it just leaves 12/18
tickets unaddressed.
tk13_share_drive(hard), C spent+22.8 % vs A ($0.02726 vs $0.02219) for the identical outcome, because
it burned two low-tier calls before landing on frontier. Full analysis in
README_PART2 §6.
96fac5e7b2e0e3dd912292062959c625, 10 spans,s15.costacross everyprovider-call span sums to $0.03257800 — exactly the ledger
spent.proofs/out/part2_trace_with_costs.json
Part 3 — same policy under an unbounded planner ($0.002 ceiling):
BudgetRefusednode failure with arefusal_logentry naming which projection failed which thresholdAll six p3 invariants pass; every refusal is attributed to principal
you/part2-attacker, reasoned, and countable in both the ledger and theOTel span export.
What is in this PR
Nothing under
s15code/is touched. The whole delta is data + config + docsHow to reproduce from a fresh checkout