Repository navigation
One daily budget for every judge call on the user's key - #842
Merged
Merged
Conversation
evaluate_with_llm_judge and verify_citations had a per-call cap and no daily limit, so an agent calling either in a loop, or steered into doing so, could spend the user's provider key without end; only the relevance judge had a daily budget. All three now draw on one, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day, per tenant), kept in the database ledger the relevance judge already used. The server builds the budget once and hands the same instance to the engine (for the judge tools) and to the relevance judge. Each judge call reserves its worst case before it is made and settles to its cost after; a refused call spends nothing. evaluate_with_llm_judge answers IRIS_BUDGET_EXCEEDED (retryable, with the reset time); verify_citations marks the citation daily_budget_reached and makes no further call. IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD is still read when the new name is unset, and the server says so at startup. Today's spend is judge.dailyBudget on iris://capabilities. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Iris gate — 1 of 2 tripped
|
| Trace | Verdict | Basis | Rules, classes or missing inputs | Evidence |
|---|---|---|---|---|
0bf1676b1d41a6a4cdb7ce60c77e1d63 |
failed | detector_veto + risk_over_loss |
no_pii, pii_leak, credential_leak | no_pii: AWS Access Key (output 45–65) |
| Verdict basis | Traces |
|---|---|
detector_veto |
2 |
clean |
1 |
Unjudged questions: task_completed (3), tool_use_correct (3) — a trace that did not carry what a rule needs.
tests/fixtures/ci-gate/traces.ndjson · 3 evaluated · dataset release-gate: 2 in the gate · exit 1 · what the bases mean
Iris gate — 1 stored, nothing tripped
|
| Verdict basis | Traces |
|---|---|
clean |
1 |
Unjudged questions: task_completed (1), tool_use_correct (1) — a trace that did not carry what a rule needs.
tests/fixtures/ci-gate/clean.ndjson · 1 evaluated · exit 0 · what the bases mean
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
evaluate_with_llm_judgeandverify_citationshad a per-call cap and no daily limit. An agent calling either in a loop, or steered into doing so by text it read, could spend the user's provider key without end. Only the relevance judge had a daily budget.All three now draw on one budget, because one key pays for all three:
IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD, relevance judge onlyIRIS_LLM_JUDGE_DAILY_BUDGET_USD, every judge call (default 1 USD per UTC day, per tenant)evaluate_with_llm_judgepast the budgetIRIS_BUDGET_EXCEEDED,retryable: true, the reset time in the message; nothing spentverify_citationspast the budgetjudge_error.kind: "daily_budget_reached"and no further call is made; when nothing was judged,IRIS_JUDGE_FAILEDnaming that kindjudge.relevance.budget, only with a relevance judge installedjudge.dailyBudgetoniris://capabilities, alwaysHow it works: the server builds the budget once over the database ledger the relevance judge already used (migration 017) and gives the same instance to the engine, for the judge tools, and to the relevance judge. Each judge call reserves its worst case (the estimate the per-call cap uses) before it is made, then settles to its actual cost; a call refused before any spend releases its reservation, and a call that failed after the provider may have billed it keeps its worst case counted.
Breaking
IRIS_LLM_JUDGE_DAILY_BUDGET_USD.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USDis still read when the new name is unset, and the server says so at startup.judge_daily_budget_usd); a value set under the old one is not carried over, so it falls back to the 1 USD default.Tests
tests/unit/tools/judge-daily-budget.test.ts: through the real tools over an in-memory MCP transport. A spent budget refusesevaluate_with_llm_judgeand every citation judge call before any provider call; both tools draw on one balance, settled to what the calls cost; the relevance judge holds the same budget instance as the tools; the old variable name still sets it.tests/unit/eval/llm-judge/judge-budget.test.ts: the old name is read when the new one is unset, gives way to it silently, and a bad value under it is named by it; the gate holds, settles, releases and refuses.npm run preflightpassed on this commit.🤖 Generated with Claude Code