Skip to content

One daily budget for every judge call on the user's key - #842

Merged
irparent merged 2 commits into
mainfrom
fix/one-daily-judge-budget
Oct 5, 2026
Merged

irparent merged 2 commits into
mainfrom
fix/one-daily-judge-budget

Conversation

@irparent

@irparent irparent commented Oct 5, 2026

Copy link
Copy Markdown
Member

What changes

evaluate_with_llm_judge and verify_citations had a per-call cap and no daily limit. An agent calling either in a loop, or steered into doing so by text it read, could spend the user's provider key without end. Only the relevance judge had a daily budget.

All three now draw on one budget, because one key pays for all three:

Before Now
Variable IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD, relevance judge only IRIS_LLM_JUDGE_DAILY_BUDGET_USD, every judge call (default 1 USD per UTC day, per tenant)
evaluate_with_llm_judge past the budget no limit IRIS_BUDGET_EXCEEDED, retryable: true, the reset time in the message; nothing spent
verify_citations past the budget no limit the citation gets judge_error.kind: "daily_budget_reached" and no further call is made; when nothing was judged, IRIS_JUDGE_FAILED naming that kind
Where to see today's spend judge.relevance.budget, only with a relevance judge installed also judge.dailyBudget on iris://capabilities, always

How it works: the server builds the budget once over the database ledger the relevance judge already used (migration 017) and gives the same instance to the engine, for the judge tools, and to the relevance judge. Each judge call reserves its worst case (the estimate the per-call cap uses) before it is made, then settles to its actual cost; a call refused before any spend releases its reservation, and a call that failed after the provider may have billed it keeps its worst case counted.

Breaking

  • A deployment that spends more than 1 USD a day on the judge tools raises IRIS_LLM_JUDGE_DAILY_BUDGET_USD.
  • IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD is still read when the new name is unset, and the server says so at startup.
  • In the Claude Desktop extension the setting is now "LLM judge daily budget" (judge_daily_budget_usd); a value set under the old one is not carried over, so it falls back to the 1 USD default.

Tests

  • tests/unit/tools/judge-daily-budget.test.ts: through the real tools over an in-memory MCP transport. A spent budget refuses evaluate_with_llm_judge and every citation judge call before any provider call; both tools draw on one balance, settled to what the calls cost; the relevance judge holds the same budget instance as the tools; the old variable name still sets it.
  • tests/unit/eval/llm-judge/judge-budget.test.ts: the old name is read when the new one is unset, gives way to it silently, and a bad value under it is named by it; the gate holds, settles, releases and refuses.
  • npm run preflight passed on this commit.

🤖 Generated with Claude Code

irparent and others added 2 commits October 5, 2026 11:45
evaluate_with_llm_judge and verify_citations had a per-call cap and no
daily limit, so an agent calling either in a loop, or steered into doing
so, could spend the user's provider key without end; only the relevance
judge had a daily budget. All three now draw on one,
IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day, per tenant),
kept in the database ledger the relevance judge already used.

The server builds the budget once and hands the same instance to the
engine (for the judge tools) and to the relevance judge. Each judge call
reserves its worst case before it is made and settles to its cost after;
a refused call spends nothing. evaluate_with_llm_judge answers
IRIS_BUDGET_EXCEEDED (retryable, with the reset time); verify_citations
marks the citation daily_budget_reached and makes no further call.

IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD is still read when the new name is
unset, and the server says so at startup. Today's spend is
judge.dailyBudget on iris://capabilities.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
website Building Building Preview Oct 5, 2026 6:46pm UTC

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Iris gate — 1 of 2 tripped --fail-on detector_veto

iris-eval ingest: 3 stored, 1 tripped --fail-on detector_veto (2 of 3 evaluated in dataset "release-gate")

Trace Verdict Basis Rules, classes or missing inputs Evidence
0bf1676b1d41a6a4cdb7ce60c77e1d63 failed detector_veto + risk_over_loss no_pii, pii_leak, credential_leak no_pii: AWS Access Key (output 45–65)
Verdict basis Traces
detector_veto 2
clean 1

Unjudged questions: task_completed (3), tool_use_correct (3) — a trace that did not carry what a rule needs.

tests/fixtures/ci-gate/traces.ndjson · 3 evaluated · dataset release-gate: 2 in the gate · exit 1 · what the bases mean

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Iris gate — 1 stored, nothing tripped --fail-on any

iris-eval ingest: 1 stored, 0 tripped --fail-on any

Verdict basis Traces
clean 1

Unjudged questions: task_completed (1), tool_use_correct (1) — a trace that did not carry what a rule needs.

tests/fixtures/ci-gate/clean.ndjson · 1 evaluated · exit 0 · what the bases mean

@irparent
irparent merged commit ce00584 into main Oct 5, 2026
74 checks passed
@irparent
irparent deleted the fix/one-daily-judge-budget branch October 5, 2026 19:13

This branch was successfully deployed

1 active deployment
Preview — cc82d3fa Deployed Oct 5, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant