diff --git a/.claims.json b/.claims.json index 9b9f467b..0d238c6b 100644 --- a/.claims.json +++ b/.claims.json @@ -5141,8 +5141,8 @@ ], "version": 1 }, - "generatedAt": "2026-10-05T15:25:24.087Z", - "generatedFromCommit": "42365584", + "generatedAt": "2026-10-05T17:06:12.498Z", + "generatedFromCommit": "547c3c39", "generatorVersion": "1.0.0", "llmJudgeTemplates": { "count": 7, @@ -5152,8 +5152,8 @@ "Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the \"env\" block of the iris-eval entry in your MCP config — \"iris-eval\": { \"command\": \"npx\", \"args\": [\"-y\", \"@iris-eval/mcp-server\"], \"env\": { \"IRIS_ANTHROPIC_API_KEY\": \"sk-ant-...\" } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval.", "Restart the MCP session. A running process never sees a variable set after it started.", "Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard.", - "Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it.", - "Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why." + "Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it.", + "Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why." ], "title": "Enable the LLM judge (optional; the deterministic rules never need it)" }, @@ -9497,7 +9497,7 @@ "passed": null, "total": 51 }, - "totalCombined": 5298, + "totalCombined": 5307, "vitestDashboard": { "failed": 0, "passed": 443, @@ -9505,8 +9505,8 @@ }, "vitestRoot": { "failed": 0, - "passed": 4804, - "total": 4804 + "passed": 4813, + "total": 4813 } }, "version": { diff --git a/CHANGELOG.md b/CHANGELOG.md index 483dedab..cae3f233 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -26,6 +26,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ### Security - **BREAKING — A tool argument can no longer widen what the operator allows for fetching or spend.** Three arguments could. `verify_citations` with `allow_fetch: true` fetched the cited URLs on a server whose operator had not set `IRIS_CITATION_ALLOW_FETCH=1`; its `domain_allowlist` was added to the operator's `IRIS_CITATION_DOMAINS` instead of narrowing it; and `max_cost_usd` on `evaluate_with_llm_judge` and `max_cost_usd_total` on `verify_citations` replaced the operator's cap with any larger number. An agent's arguments can be steered by text it read, including text stored in Iris, so each is now a ceiling: an argument can narrow the operator's setting for one call and never widen it. When one asks for more, the operator's setting applies and the call goes on, the response carries a warning with code `IRIS_ARGUMENT_NARROWED` naming the argument and the setting that applied, and the stored evaluation keeps the same fact (`provenance.narrowed`) with a sentence to the operator in its `interpretations`, on every later read. Fetching now needs `IRIS_CITATION_ALLOW_FETCH=1` on the server; a caller that relied on `allow_fetch: true` gets no fetch and the warning. The per-call cap on `verify_citations` is a new setting, `IRIS_CITATION_MAX_COST_USD_TOTAL`, default 1 USD as before. +- **BREAKING — Every judge call on your key draws on one daily budget.** `evaluate_with_llm_judge` and `verify_citations` had a per-call cap and no daily limit, so an agent calling either in a loop, or steered into doing so, could spend the key without end; only the relevance judge had a daily budget. All three now share one, `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` (default 1 USD per UTC day, per tenant, kept in the database). A call is made only if its worst case fits in what is left. When it does not, nothing is spent: `evaluate_with_llm_judge` answers `IRIS_BUDGET_EXCEEDED` (retryable, with the time the budget resets), and `verify_citations` marks the citation `daily_budget_reached` and makes no further call. A deployment that spends more than 1 USD a day on the judge tools raises the budget. The variable was `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD`; that name is still read when the new one is unset, and the server says so at startup. In the Claude Desktop extension the setting is now "LLM judge daily budget", and a value set under the old one is not carried over. Today's spend is `judge.dailyBudget` on `iris://capabilities`. ## [0.20.0] - 2026-10-03 diff --git a/README.md b/README.md index 029e565e..af9a63ab 100644 --- a/README.md +++ b/README.md @@ -354,8 +354,8 @@ Iris registers twelve tools that any MCP-compatible agent can invoke — trace a 2. Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the "env" block of the iris-eval entry in your MCP config — "iris-eval": { "command": "npx", "args": ["-y", "@iris-eval/mcp-server"], "env": { "IRIS_ANTHROPIC_API_KEY": "sk-ant-..." } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval. 3. Restart the MCP session. A running process never sees a variable set after it started. 4. Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard. -5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it. -6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. +5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it. +6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. When `IRIS_OTEL_ENDPOINT` is configured, `log_trace` calls also emit a best-effort OTLP/HTTP JSON export to any OpenTelemetry collector (Jaeger, Grafana Tempo, Datadog OTLP, Honeycomb, etc). See [docs/otel-integration.md](https://github.com/iris-eval/mcp-server/blob/main/docs/otel-integration.md). @@ -485,7 +485,8 @@ Every variable `--help` documents. CLI flags take precedence over environment va | `IRIS_OPENAI_API_KEY` | Required by `evaluate_with_llm_judge` + `verify_citations` with `provider=openai` | | `IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL` | Hard cost cap per LLM judge call (default `0.25`). `max_cost_usd` can lower it for a call, never raise it | | `IRIS_RELEVANCE_JUDGE_MODEL` | A priced judge model id (e.g. `claude-haiku-4-5`). When set, with that provider's key, `answers_the_ask` asks this LLM judge on every evaluation that carries an input and gates on its relevance verdict — one judge call per evaluation, under the cost cap above and the two limits below. **Each such evaluation's input and output are sent to that model's provider (Anthropic or OpenAI) on your key**, with the personal data and credentials `no_pii` flags replaced first. Unset (the default), `answers_the_ask` reads the ask lexically and advises, and nothing is sent ([docs/llm-as-judge.md](https://github.com/iris-eval/mcp-server/blob/main/docs/llm-as-judge.md#the-relevance-judge-behind-answers_the_ask)) | -| `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD` | What the relevance judge may spend per UTC day, per tenant (default `1`). Kept in the database, so a restart does not reset it. A call is made only if its worst case fits in what is left; past that, `answers_the_ask` reads the ask lexically and `judge.withheld` is `daily_budget`. `0` stops every call | +| `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` | What every judge call together may spend per UTC day, per tenant (default `1`): the relevance judge, `evaluate_with_llm_judge` and `verify_citations` share it. Kept in the database, so a restart does not reset it. A call is made only if its worst case fits in what is left; past that, `answers_the_ask` reads the ask lexically (`judge.withheld: "daily_budget"`) and the judge tools answer `IRIS_BUDGET_EXCEEDED` (a citation, `daily_budget_reached`). `0` stops every call | +| `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD` | The name of `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` before 0.21.0, when it limited the relevance judge only. Read when the new name is unset, and it then limits every judge call; the server says so at startup | | `IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST` | Relevance judge calls one request may make (default `20`): an OTLP batch or an `evaluate_runs` re-score judges its first 20 traces and reads the rest lexically, with `judge.withheld: "request_cap"` | | `IRIS_RELEVANCE_JUDGE_REDACT` | `on` (default): every span `no_pii` flags (personal data and credentials) in the input and output is replaced by a `[REDACTED:#]` marker before they are sent to the relevance judge. `off` sends them as they are | | `IRIS_CITATION_ALLOW_FETCH` | Set to `1` to permit outbound HTTP in `verify_citations` (off by default). A tool argument can turn it off for a call, never on | diff --git a/claude-plugin/skills/iris-eval/SKILL.md b/claude-plugin/skills/iris-eval/SKILL.md index af2b15bb..150732df 100644 --- a/claude-plugin/skills/iris-eval/SKILL.md +++ b/claude-plugin/skills/iris-eval/SKILL.md @@ -237,8 +237,8 @@ result, including a passing one. 2. Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the "env" block of the iris-eval entry in your MCP config — "iris-eval": { "command": "npx", "args": ["-y", "@iris-eval/mcp-server"], "env": { "IRIS_ANTHROPIC_API_KEY": "sk-ant-..." } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval. 3. Restart the MCP session. A running process never sees a variable set after it started. 4. Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard. -5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it. -6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. +5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it. +6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. ## Example Workflows diff --git a/docs/api-reference.md b/docs/api-reference.md index 6815a488..80457c03 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -861,12 +861,14 @@ Each citation carries at most one of two error fields, one per stage: | Field | Set when | Kinds | |-------|----------|-------| | `resolve_error` | The source was not resolved (`resolve_status` is `skipped` or `error`) | `unresolvable_kind`, `fetch_disabled`, `bad_scheme`, `ssrf`, `not_allowed_domain`, `timeout`, `too_large`, `bad_status`, `redirect_loop`, `not_text` | -| `judge_error` | The source resolved (`resolve_status` is `ok`) and the judge gave no verdict, so `judge` is absent | `cost_cap_reached`, `malformed_judge_response`, and the provider error kinds `auth`, `rate_limit`, `bad_request`, `server_error`, `timeout`, `malformed_response`, `unknown` | +| `judge_error` | The source resolved (`resolve_status` is `ok`) and the judge gave no verdict, so `judge` is absent | `cost_cap_reached`, `daily_budget_reached`, `malformed_judge_response`, and the provider error kinds `auth`, `rate_limit`, `bad_request`, `server_error`, `timeout`, `malformed_response`, `unknown` | A citation with either error is unverified, never unsupported: it is left out of `total_judged` and of the score. Until 0.20.0 judge failures were reported under `resolve_error` on citations whose `resolve_status` was `ok`. **The operator's settings are ceilings.** Whether sources are fetched, from which domains, and what a call may spend are the operator's settings: `IRIS_CITATION_ALLOW_FETCH`, `IRIS_CITATION_DOMAINS` and `IRIS_CITATION_MAX_COST_USD_TOTAL` here, `IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL` for `evaluate_with_llm_judge`. An argument can narrow them for one call and never widen them, because an agent's arguments can be steered by text it read. When an argument asks for more, the operator's setting applies and the call goes on: `warnings` carries an entry with code `IRIS_ARGUMENT_NARROWED`, the argument in `field`, and a sentence naming what applied and the setting that would allow more. The stored evaluation keeps the same fact in `provenance.narrowed`, with an `interpretations` sentence to the operator, so it is there on every later read. Until 0.21.0, `allow_fetch: true` turned fetching on without the operator's setting, `domain_allowlist` was added to `IRIS_CITATION_DOMAINS`, and either cost argument could replace the operator's cap with a larger one. +**One daily budget.** Every judge call Iris makes on the key, from `evaluate_with_llm_judge`, `verify_citations` and the relevance judge, draws on the operator's `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` (default 1 USD per UTC day, per tenant, kept in the database). A call is made only if its worst case fits in what is left. Refused, `evaluate_with_llm_judge` answers `IRIS_BUDGET_EXCEEDED` (`retryable: true`, the reset time in the message) and nothing is spent; `verify_citations` marks the citation `daily_budget_reached` and makes no further call, and when no citation was judged it answers `IRIS_JUDGE_FAILED` naming that kind. Today's spend is `judge.dailyBudget` on `iris://capabilities`. + **SSRF defense + auth:** Eight layers documented in [semantic-citation-verify.md](./semantic-citation-verify.md). Requires an LLM judge API key (same as `evaluate_with_llm_judge`). --- @@ -908,7 +910,7 @@ MCP resources are read-only data endpoints accessed via the MCP `resources/read` ### iris://capabilities -What this server can do, as one object — the same one `GET /api/v1/capabilities` serves: `version`, `transport`, the evaluation `questions` registry, the `rules` roster (each with `kind`, `mechanism`, `needs`, `question`, `classes`, `version`, the effective `critical` flag with `criticalSource`, and `proof`), `customRules` counts, the `judge` state (`enabled`, `provider`, `providers`, `costCapUsd`, `howToEnable[]` — provider name only, never a key), the `citations` posture (`fetchAllowed`, `domainsRestricted`), the `dashboard` address and mode, the `limits` a caller will hit, the `tools`, `resources` and `prompts` registered, and `toolGuide`: per tool, the long form its capped description points to (`does`, `whenNot`, `errors`, the full explanation of some `parameters`, and what each output field `returns`). +What this server can do, as one object — the same one `GET /api/v1/capabilities` serves: `version`, `transport`, the evaluation `questions` registry, the `rules` roster (each with `kind`, `mechanism`, `needs`, `question`, `classes`, `version`, the effective `critical` flag with `criticalSource`, and `proof`), `customRules` counts, the `judge` state (`enabled`, `provider`, `providers`, `costCapUsd`, `howToEnable[]` — provider name only, never a key — the `relevance` judge, and `dailyBudget`: today's spend against `IRIS_LLM_JUDGE_DAILY_BUDGET_USD`, which every judge call draws on), the `citations` posture (`fetchAllowed`, `domainsRestricted`, and `totalCostCapUsd`, the most one `verify_citations` call may spend), the `dashboard` address and mode, the `limits` a caller will hit, the `tools`, `resources` and `prompts` registered, and `toolGuide`: per tool, the long form its capped description points to (`does`, `whenNot`, `errors`, the full explanation of some `parameters`, and what each output field `returns`). ### iris://proof diff --git a/docs/llm-as-judge.md b/docs/llm-as-judge.md index 9a4b07cf..4f16310d 100644 --- a/docs/llm-as-judge.md +++ b/docs/llm-as-judge.md @@ -55,8 +55,8 @@ await callTool('evaluate_with_llm_judge', { 2. Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the "env" block of the iris-eval entry in your MCP config — "iris-eval": { "command": "npx", "args": ["-y", "@iris-eval/mcp-server"], "env": { "IRIS_ANTHROPIC_API_KEY": "sk-ant-..." } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval. 3. Restart the MCP session. A running process never sees a variable set after it started. 4. Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard. -5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it. -6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. +5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it. +6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. ### 2. Optional: set a stricter cost cap @@ -114,17 +114,17 @@ Set `IRIS_RELEVANCE_JUDGE_MODEL` to a priced model id, with that provider's key, - Why it is the default: relevance is about subject, not about the digits of a card number, so replacing a value with a labelled marker leaves the question the judge answers intact. Sending a secret Iris itself flags to a third party would be the leak the product exists to catch. The cost is an evaluation whose whole subject is a value the detector flags (an ask to "repeat the key back", say): the judge sees a marker, not the key. - The opt-out: `IRIS_RELEVANCE_JUDGE_REDACT=off` sends both texts as they are, and every judge record then carries `sentUnredacted: true`. Any other value keeps redaction on. - **Spend limits.** The per-call cap bounds one judgment. Two more limits bound the total, and when either stops a call nothing is spent, the rule reads the ask lexically (it advises), and `judge.withheld` says which limit it was, with the sentence in `judge.error`: - - `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD` (default `1`) is what the judge may spend per UTC day, per tenant. It is kept in the database (table `relevance_judge_spend`, migration 017), so a restart, a second server on the same file and `iris-eval ingest` all draw on one balance. A call is made only if its worst case (the same two-attempt estimate the per-call cap uses) fits in what is left; that worst case is held while the call runs and replaced by the call's actual cost when it returns, so the day's total never passes the budget. A call that fails after the provider may have billed it stays counted at its worst case. `withheld: "daily_budget"`, and the first refusal of the day for a tenant writes one warning line to the server log (once per process). `0` stops every call. The budget resets at 00:00 UTC, the day the providers' usage pages report by. At the $0.0060 worst case above, the default admits at least 165 judgments a day, and more once calls settle to what they cost. + - `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` (default `1`) is what every judge call together may spend per UTC day, per tenant: this judge's calls and those of `evaluate_with_llm_judge` and `verify_citations` draw on one balance, because one key pays for all three. It is kept in the database (table `relevance_judge_spend`, migration 017), so a restart, a second server on the same file and `iris-eval ingest` all draw on it too. Until 0.21.0 it was `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD` and covered this judge only; that name is still read when the new one is unset, and the server says so at startup. A call is made only if its worst case (the same two-attempt estimate the per-call cap uses) fits in what is left; that worst case is held while the call runs and replaced by the call's actual cost when it returns, so the day's total never passes the budget. A call that fails after the provider may have billed it stays counted at its worst case. `withheld: "daily_budget"`, and the first refusal of the day for a tenant writes one warning line to the server log (once per process). `0` stops every call. The budget resets at 00:00 UTC, the day the providers' usage pages report by. At the $0.0060 worst case above, the default admits at least 165 judgments a day, and more once calls settle to what they cost. - `IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST` (default `20`) is how many judge calls one request may make. An OTLP request can carry 2,000 traces and `evaluate_runs` re-scores a whole run; each judges its first 20 and reads the rest lexically with `withheld: "request_cap"`. The OTLP answer's `iris-eval.relevance_judge` gives `calls`, `withheld` and `max_calls_per_request`, and the `evaluate_runs` summary says how many were withheld. `iris-eval ingest` is not one request: each trace it reads is its own, and the daily budget bounds a large file. - **A key alone never turns it on.** The key enables `evaluate_with_llm_judge`, which you call and pay for per call. The model has to be named because cost varies a hundredfold across models, and so that a key set for the judge tool does not start billing every evaluation. - **What the result carries.** `rule_results[answers_the_ask]` has `kind: "judgment"`, `role: "gate"`, the judge's `sample` evidence, and a `judge` object: `provider`, `model`, `score`, `passThreshold` (0.60), `passed`, `rationale`, `dimensions`, `costUsd`, tokens and latency. The verdict is the score against the pass line; the model's own `passed` is kept as `selfReportedPass` and never obeyed. - **What it judges that the words cannot.** A paraphrase that shares none of the ask's words, a one-word answer, and an ask with a single content term are all judged, not skipped. A wrong answer to the question asked is relevant; correctness is another template's question. - **When the judge cannot answer.** An unpriced model, a missing key, a cost-cap refusal, a spend limit or a provider error is recorded as `judge.error`. The rule then falls back to the lexical reading, which advises, and `interpretations` carries a warning naming the reason. A deployment that believes the judge is on never reads a lexical verdict as the judge's. A judge that is configured and cannot run at all (no key for its provider, or a model with no price) is also said at startup: the server and `iris-eval ingest` print one warning line naming what is missing, and `--self-test` fails its judge step instead of printing PASS, because every evaluation would otherwise fall back to a reading that only advises. - **Same family.** When the linked trace records the agent's model (`metadata.model` or a span's `gen_ai.request.model`) and the judge shares its family, the verdict stands and `interpretations` carries the same-family warning `evaluate_with_llm_judge` returns. -- **Where to check.** `iris://capabilities` → `judge.relevance` says whether a relevance judge is configured and whether it can be called, what it sends (`egress`, `redact`), the per-request cap, and today's `budget` (`limitUsd`, `spentUsd`, `remainingUsd`, `calls`, `refused`, `exhausted`, `resetsAt`). `--self-test` prints the same line, with today's spend read from the install's database. `GET /api/v1/health` (and `/health` on the HTTP transport) carries `judge.relevance.budget_exhausted` and `budget_resets_at`, so a probe can alert when the judge stops for the day; it does not report the amount, since it answers without a key. +- **Where to check.** `iris://capabilities` → `judge.relevance` says whether a relevance judge is configured and whether it can be called, what it sends (`egress`, `redact`), the per-request cap, and today's `budget` (`limitUsd`, `spentUsd`, `remainingUsd`, `calls`, `refused`, `exhausted`, `resetsAt`); `judge.dailyBudget` is the same balance, there whether or not a relevance judge is installed, since the judge tools draw on it too. `--self-test` prints the same line, with today's spend read from the install's database. `GET /api/v1/health` (and `/health` on the HTTP transport) carries `judge.relevance.budget_exhausted` and `budget_resets_at`, so a probe can alert when the judge stops for the day; it does not report the amount, since it answers without a key. - **Tool annotations.** With a relevance judge installed, `evaluate_output`, `log_trace` and `evaluate_runs` advertise `openWorldHint: true`, and `evaluate_output` advertises `idempotentHint: false`: each call may reach the provider and spend again. Without one they stay closed-world. Annotations are read once, when the server registers its tools, which is after it installs the judge from the environment. - **Replays.** A judge changes the ruleset hash, so `evaluate_runs` re-scores traces that were judged without it. -- **Embedders.** The engine calls no model unless you install one: `engine.setRelevanceJudge(createRelevanceJudge({ model, apiKey }))`, both exported from `@iris-eval/mcp-server/engine`. Redaction is on and the per-request cap applies there too. Without a `budget` option the daily budget is kept in the process's memory and resets when it exits; pass a `JudgeBudget` over your own ledger to keep it. +- **Embedders.** The engine calls no model unless you install one: `engine.setRelevanceJudge(createRelevanceJudge({ model, apiKey }))`, both exported from `@iris-eval/mcp-server/engine`. Redaction is on and the per-request cap applies there too. Without a `budget` option the daily budget is kept in the process's memory and resets when it exits; pass a `JudgeBudget` over your own ledger to keep it, and the same one to `engine.setJudgeBudget` so every judge call shares it. How accurate the judge is, and how the judged rule compares with the lexical one on the rule's own 45 labelled cases, is measured by `npm run proof:judge` (see [the measurement](#measuring-the-relevance-judge) below). Until a keyed run is published, those numbers are pending. diff --git a/docs/semantic-citation-verify.md b/docs/semantic-citation-verify.md index 7916ea4b..9aee4647 100644 --- a/docs/semantic-citation-verify.md +++ b/docs/semantic-citation-verify.md @@ -126,6 +126,7 @@ Same structure as `evaluate_with_llm_judge`: - **Per-call cost cap** — `max_cost_usd_total`, at most the operator's `IRIS_CITATION_MAX_COST_USD_TOTAL` (default $1.00): an argument can lower it for a call, never raise it. Budget across all judge calls in one `verify_citations` invocation. When the next call's pessimistic estimate would push total cost past the cap, the pipeline stops: the citation it stopped on reports `judge_error.kind: "cost_cap_reached"`, and later citations are not attempted (`total_citations_found` still counts them). - **Per-citation pessimistic estimate** — before each judge call, worst-case cost is computed. If adding it exceeds the cap, skip. +- **Daily budget** — every judge call Iris makes on the key, this tool's, `evaluate_with_llm_judge`'s and the relevance judge's, draws on the operator's `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` (default $1.00 per UTC day, per tenant, kept in the database). Each citation's worst case must fit in what is left today; one that does not gets `judge_error.kind: "daily_budget_reached"`, and later citations are not attempted. - **Typical cost** — on haiku with 5 citations averaging 2KB source text: ~$0.002-$0.005 total. On opus with the same: ~$0.10-$0.25. --- @@ -171,6 +172,7 @@ Set when `resolve_status` is `ok` and `judge` is absent. Until 0.20.0 these were | `judge_error.kind` | Meaning | |---------------------------|---------------------------------------------------------------------------| | `cost_cap_reached` | The next judge call would pass `max_cost_usd_total`; later citations are not attempted | +| `daily_budget_reached` | The next judge call does not fit in what is left of the day's `IRIS_LLM_JUDGE_DAILY_BUDGET_USD`; later citations are not attempted | | `malformed_judge_response`| The judge replied, but not with a readable verdict | | `auth` | The provider refused the key | | `rate_limit` | The provider rate-limited the call, after one retry | diff --git a/mcpb/manifest.json b/mcpb/manifest.json index 810cd9db..dd1cc120 100644 --- a/mcpb/manifest.json +++ b/mcpb/manifest.json @@ -33,7 +33,7 @@ "IRIS_ANTHROPIC_API_KEY": "${user_config.anthropic_api_key}", "IRIS_OPENAI_API_KEY": "${user_config.openai_api_key}", "IRIS_RELEVANCE_JUDGE_MODEL": "${user_config.relevance_judge_model}", - "IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD": "${user_config.relevance_judge_daily_budget_usd}" + "IRIS_LLM_JUDGE_DAILY_BUDGET_USD": "${user_config.judge_daily_budget_usd}" } } }, @@ -138,10 +138,10 @@ "required": false, "default": "" }, - "relevance_judge_daily_budget_usd": { + "judge_daily_budget_usd": { "type": "number", - "title": "Relevance judge daily budget (USD)", - "description": "The most the relevance judge may spend per UTC day. Past it, answers_the_ask reads the ask lexically until midnight UTC, and each result says so. 0 stops every call.", + "title": "LLM judge daily budget (USD)", + "description": "The most every judge call together may spend per UTC day: the relevance judge, evaluate_with_llm_judge and verify_citations. Past it, no judge call is made until midnight UTC, and each result says so. 0 stops every call.", "required": false, "default": 1, "min": 0 diff --git a/server.json b/server.json index 678413fc..13d20942 100644 --- a/server.json +++ b/server.json @@ -126,7 +126,14 @@ "name": "IRIS_RELEVANCE_JUDGE_MODEL" }, { - "description": "What the relevance judge may spend per UTC day, per tenant, in USD (default 1); past it answers_the_ask reads the ask lexically", + "description": "What every judge call together may spend per UTC day, per tenant, in USD (default 1): the relevance judge, evaluate_with_llm_judge and verify_citations. Past it no judge call is made", + "isRequired": false, + "format": "string", + "isSecret": false, + "name": "IRIS_LLM_JUDGE_DAILY_BUDGET_USD" + }, + { + "description": "The name of IRIS_LLM_JUDGE_DAILY_BUDGET_USD before 0.21.0. Read when the new name is unset, and it then limits every judge call", "isRequired": false, "format": "string", "isSecret": false, diff --git a/skills/iris-eval/SKILL.md b/skills/iris-eval/SKILL.md index c49bf564..856d7f2b 100644 --- a/skills/iris-eval/SKILL.md +++ b/skills/iris-eval/SKILL.md @@ -241,8 +241,8 @@ result, including a passing one. 2. Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the "env" block of the iris-eval entry in your MCP config — "iris-eval": { "command": "npx", "args": ["-y", "@iris-eval/mcp-server"], "env": { "IRIS_ANTHROPIC_API_KEY": "sk-ant-..." } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval. 3. Restart the MCP session. A running process never sees a variable set after it started. 4. Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard. -5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it. -6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. +5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it. +6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why. ## Example Workflows diff --git a/src/capabilities.ts b/src/capabilities.ts index cff84a29..67db2b1e 100644 --- a/src/capabilities.ts +++ b/src/capabilities.ts @@ -21,6 +21,7 @@ import { publishedAccuracyFor, publishedProvenance, ppvAt, type PublishedRuleAcc import { REGEX_MATCH_BUDGET_MS } from './eval/rules/regex-sandbox.js'; import { JUDGE_ENABLE_STEPS, judgeState, type JudgeProvider } from './judge-enablement.js'; import { relevanceJudgeState, type RelevanceJudgeState } from './eval/llm-judge/relevance-judge.js'; +import type { BudgetToday } from './eval/llm-judge/budget.js'; import { LOCAL_TENANT } from './types/tenant.js'; import { citationTotalCostCapUsd } from './tools/operator-ceilings.js'; import { TOOL_NAMES } from './tools/index.js'; @@ -56,6 +57,8 @@ export interface Capabilities { howToEnable: readonly string[]; /** The relevance judge answers_the_ask gates on (#649): installed only when IRIS_RELEVANCE_JUDGE_MODEL names a model. */ relevance: RelevanceJudgeState; + /** Today's spend against IRIS_LLM_JUDGE_DAILY_BUDGET_USD, which every judge call draws on: the relevance judge's and the judge tools'. */ + dailyBudget: BudgetToday | null; }; /** totalCostCapUsd: the most one verify_citations call may spend (IRIS_CITATION_MAX_COST_USD_TOTAL); an argument can lower it, never raise it. */ citations: { fetchAllowed: boolean; domainsRestricted: boolean; totalCostCapUsd: number }; @@ -116,6 +119,7 @@ export function buildCapabilities(ctx: CapabilitiesContext): Capabilities { costCapUsd: judge.costCapUsd, howToEnable: JUDGE_ENABLE_STEPS, relevance: relevanceJudgeState(ctx.evalEngine?.relevanceJudgeInForce() ?? null), + dailyBudget: ctx.evalEngine?.judgeBudgetInForce()?.today(LOCAL_TENANT) ?? null, }, citations: { fetchAllowed: process.env.IRIS_CITATION_ALLOW_FETCH === '1', diff --git a/src/eval/citation-verify/verifier.ts b/src/eval/citation-verify/verifier.ts index feb864b4..8179b8e2 100644 --- a/src/eval/citation-verify/verifier.ts +++ b/src/eval/citation-verify/verifier.ts @@ -11,6 +11,7 @@ import { SECURITY_NOTICE, TAIL_REINFORCEMENT, } from '../llm-judge/templates/index.js'; +import type { JudgeSpendGate } from '../llm-judge/budget.js'; import { extractCitations, type ExtractedCitation } from './extract.js'; import { resolveSource, CitationResolveError, type ResolvedSource } from './resolve.js'; @@ -26,6 +27,13 @@ export interface VerifyCitationsParams { perSourceMaxBytes?: number; // Cap number of citations we attempt — protects against DoS-by-spam. maxCitations?: number; + /** + * The daily budget every judge call on this key draws on, when the caller + * keeps one (llm-judge/budget.ts): each call's worst case is held while it + * runs and replaced by its cost. A citation the budget refuses gets + * judgeError `daily_budget_reached`, and no further call is made. + */ + spend?: JudgeSpendGate; } export interface VerifiedCitation { @@ -284,6 +292,22 @@ export async function verifyCitations( }); break; // No point continuing — subsequent calls will also exceed. } + const hold = params.spend?.(pessimistic) ?? null; + if (hold !== null && !hold.ok) { + out.push({ + citation, + resolveStatus: 'ok', + source: { + url: source.url, + status: source.status, + contentType: source.contentType, + bytesFetched: source.bytesFetched, + truncated: source.truncated, + }, + judgeError: { kind: 'daily_budget_reached', message: hold.reason }, + }); + break; // The budget is per day: every later call would be refused too. + } let judgeResponse; try { @@ -297,6 +321,8 @@ export async function verifyCitations( apiKey: params.apiKey, }); } catch (err) { + // The provider may have billed a call that failed: its worst case stays counted. + if (hold?.ok) hold.settle(null); const e = err as Error; out.push({ citation, @@ -318,6 +344,7 @@ export async function verifyCitations( const cost = estimateCostUsd(params.model, judgeResponse.inputTokens, judgeResponse.outputTokens); totalCost += cost ?? 0; + if (hold?.ok) hold.settle(cost); let parsed; try { diff --git a/src/eval/engine.ts b/src/eval/engine.ts index 272959f9..622fea26 100644 --- a/src/eval/engine.ts +++ b/src/eval/engine.ts @@ -25,6 +25,7 @@ import { PUBLISHED_CALIBRATION } from './published-calibration.js'; import { answersTheAsk } from './rules/relevance.js'; import { agentModelOf } from './llm-judge/family.js'; import type { RelevanceJudge } from './llm-judge/relevance-judge.js'; +import type { JudgeBudget } from './llm-judge/budget.js'; import { PKG_VERSION } from '../config/defaults.js'; import { generateEvalId } from '../utils/ids.js'; @@ -146,6 +147,14 @@ export class EvalEngine { */ private relevanceJudge: RelevanceJudge | null = null; + /** + * The daily budget every judge call on the user's key draws on: the + * relevance judge's and the judge tools' (llm-judge/budget.ts). The + * server sets it from the environment over the database's ledger and + * hands the same one to the relevance judge, so the three share a balance. + */ + private judgeBudget: JudgeBudget | null = null; + /** * One structured event per evaluation, whichever door asked for it * (0.15.0): the server wires it to the logger's `event('evaluation', …)`. @@ -173,6 +182,16 @@ export class EvalEngine { return this.relevanceJudge; } + /** Set the daily budget the judge tools draw on; pass the same one to the relevance judge. */ + setJudgeBudget(budget: JudgeBudget | null): void { + this.judgeBudget = budget; + } + + /** The daily budget in force: the one set, else the relevance judge's, else none. */ + judgeBudgetInForce(): JudgeBudget | null { + return this.judgeBudget ?? this.relevanceJudge?.budget ?? null; + } + /** The judge's identity as the ruleset and config hashes fold it in: absent without a judge, so no existing hash moves. */ private judgeIdentity(): string | undefined { return this.relevanceJudge ? `relevance:${this.relevanceJudge.provider ?? 'unknown'}/${this.relevanceJudge.model}` : undefined; diff --git a/src/eval/llm-judge/budget.ts b/src/eval/llm-judge/budget.ts index e24bf9ff..812d1a3b 100644 --- a/src/eval/llm-judge/budget.ts +++ b/src/eval/llm-judge/budget.ts @@ -1,24 +1,31 @@ /* - * The relevance judge's spend guards. + * The judge's spend guards. * - * The relevance judge calls a provider on the user's own key once per - * evaluation that carries an input, and evaluations arrive from every door: - * evaluate_output, log_trace, an OTLP batch of up to 2,000 traces, an - * evaluate_runs re-score of a whole run. The per-call cap - * (IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL) bounds one call; nothing bounded - * the sum. Two limits do now, and each one, when it stops a call, leaves - * the evaluation to answers_the_ask's lexical reading with the reason on the - * result — a verdict is never withheld for money, only the judge is: + * Iris calls a provider on the user's own key from three places: the + * relevance judge, once per evaluation that carries an input (from every + * door: evaluate_output, log_trace, an OTLP batch of up to 2,000 traces, an + * evaluate_runs re-score of a whole run), and the two judge tools a caller + * invokes, evaluate_with_llm_judge and verify_citations. The per-call caps + * (IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL, IRIS_CITATION_MAX_COST_USD_TOTAL) + * bound one call; the limits here bound the sum. When one stops a call, the + * call is not made and the reason is on the result: the relevance judge + * leaves the evaluation to answers_the_ask's lexical reading (a verdict is + * never withheld for money, only the judge is), and a judge tool answers + * IRIS_BUDGET_EXCEEDED, or for a citation `daily_budget_reached`. * - * IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD what the judge may spend per - * UTC day, per tenant. Counted in the database beside the traces, so a - * restart, a second server or a CLI ingest on the same file all draw on - * one balance. A call is admitted only when its WORST case (the same - * pessimistic two-attempt estimate the per-call cap uses) fits in what - * is left, then the reservation is replaced by the provider's actual - * cost — so the budget is a ceiling, never a target to overshoot. - * IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST judge calls one request may - * make, so one batch cannot spend the day's budget in one go. + * IRIS_LLM_JUDGE_DAILY_BUDGET_USD what every judge call together may + * spend per UTC day, per tenant. One balance, because one key pays: an + * agent steered into calling a judge tool in a loop spends the same + * dollars the relevance judge does. Counted in the database beside the + * traces, so a restart, a second server or a CLI ingest on the same + * file all draw on it. A call is admitted only when its WORST case (the + * same pessimistic estimate the per-call cap uses) fits in what is left, + * then the reservation is replaced by the provider's actual cost — so + * the budget is a ceiling, never a target to overshoot. Until 0.21.0 it + * was IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD and covered the relevance + * judge only; that name is still read when the new one is unset. + * IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST relevance judge calls one + * request may make, so one batch cannot spend the day's budget in one go. * * UTC days rather than a rolling window: the providers' usage pages report * by UTC day, so the number Iris shows is the number the bill shows, and @@ -26,17 +33,20 @@ */ import type { TenantId } from '../../types/tenant.js'; -export const DAILY_BUDGET_VAR = 'IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD'; +export const DAILY_BUDGET_VAR = 'IRIS_LLM_JUDGE_DAILY_BUDGET_USD'; +/** The name before 0.21.0, when the budget covered the relevance judge only. Read when DAILY_BUDGET_VAR is unset. */ +export const LEGACY_DAILY_BUDGET_VAR = 'IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD'; export const MAX_CALLS_PER_REQUEST_VAR = 'IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST'; /** - * 1 USD a day. Low on purpose: the judge is opted into by naming a model, - * and a user who did that for a handful of evaluations must not find a - * surprise on their bill because an SDK started sending every call. The - * worst case of one relevance judgment on claude-haiku-4-5 is $0.0060 on - * the proof cases (tests/unit/proof/relevance-judge-proof.test.ts), so a - * dollar admits at least 165 judgments a day before any reservation is - * settled down to its actual cost; a deployment that wants more raises it. + * 1 USD a day. Low on purpose: a user who named a relevance judge model + * for a handful of evaluations must not find a surprise on their bill + * because an SDK started sending every call, and an agent steered into + * calling a judge tool in a loop must not leave one either. The worst case + * of one relevance judgment on claude-haiku-4-5 is $0.0060 on the proof + * cases (tests/unit/proof/relevance-judge-proof.test.ts), so a dollar + * admits at least 165 judgments a day before any reservation is settled + * down to its actual cost; a deployment that wants more raises it. */ export const DEFAULT_DAILY_BUDGET_USD = 1; @@ -59,16 +69,30 @@ export interface Setting { note?: string; } -/** The daily budget from the environment: any number ≥ 0 (0 stops every judge call), else the default. */ +/** + * The daily budget from the environment: any number ≥ 0 (0 stops every + * judge call), else the default. IRIS_LLM_JUDGE_DAILY_BUDGET_USD, or when + * that is unset its old name, which now limits every judge call too. + */ export function dailyBudgetUsd(): Setting { - const raw = process.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD?.trim(); + const current = process.env.IRIS_LLM_JUDGE_DAILY_BUDGET_USD?.trim(); + const legacy = process.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD?.trim(); + const name = current ? DAILY_BUDGET_VAR : LEGACY_DAILY_BUDGET_VAR; + const raw = current || legacy; if (!raw) return { value: DEFAULT_DAILY_BUDGET_USD, source: 'default' }; const n = Number(raw); - if (Number.isFinite(n) && n >= 0) return { value: n, source: 'env' }; + if (!Number.isFinite(n) || n < 0) { + return { + value: DEFAULT_DAILY_BUDGET_USD, + source: 'default', + note: `${name}="${raw.slice(0, 40)}" is not a number of dollars ≥ 0, so the default ${DEFAULT_DAILY_BUDGET_USD} USD applies`, + }; + } + if (current) return { value: n, source: 'env' }; return { - value: DEFAULT_DAILY_BUDGET_USD, - source: 'default', - note: `${DAILY_BUDGET_VAR}="${raw.slice(0, 40)}" is not a number of dollars ≥ 0, so the default ${DEFAULT_DAILY_BUDGET_USD} USD applies`, + value: n, + source: 'env', + note: `${LEGACY_DAILY_BUDGET_VAR} is the old name of ${DAILY_BUDGET_VAR}; its ${n} USD now limits every judge call, the judge tools included`, }; } @@ -169,6 +193,17 @@ export interface BudgetToday { resetsAt: string; } +/** + * Ask to spend up to `worstCaseUsd` on one judge call. Granted: make the + * call, then `settle` with its actual cost (null when it failed after the + * provider may have billed it, which keeps the worst case counted), or + * `release` when it was refused before any spend. Refused: do not make the + * call; `reason` says why, for the result. + */ +export type JudgeSpendGate = ( + worstCaseUsd: number, +) => { ok: true; settle(actualUsd: number | null): void; release(): void } | { ok: false; reason: string }; + export interface JudgeBudgetOptions { dailyUsd: number; ledger: JudgeSpendLedger; @@ -209,15 +244,15 @@ export class JudgeBudget { const spent = this.ledger.read(tenantId, day).spentMicroUsd; const left = Math.max(0, this.limitMicroUsd - spent); const reason = - `the relevance judge's daily budget has ${fromMicroUsd(left).toFixed(4)} of ${this.dailyUsd} USD left today (UTC), ` + + `the daily judge budget has ${fromMicroUsd(left).toFixed(4)} of ${this.dailyUsd} USD left today (UTC), ` + `and this call could cost up to ${worstCaseUsd.toFixed(4)} USD, so it was not made; the budget resets at ${nextUtcMidnight(now)} ` + `(${DAILY_BUDGET_VAR} raises it)`; const key = `${tenantId}\u0000${day}`; if (!this.announced.has(key)) { this.announced.add(key); this.log( - `Relevance judge: the daily budget (${this.dailyUsd} USD, ${DAILY_BUDGET_VAR}) is spent for tenant ${tenantId} on ${day} UTC; ` + - `answers_the_ask reads the ask lexically until ${nextUtcMidnight(now)}, and each result says so.`, + `LLM judge: the daily budget (${this.dailyUsd} USD, ${DAILY_BUDGET_VAR}) is spent for tenant ${tenantId} on ${day} UTC; ` + + `until ${nextUtcMidnight(now)} no judge call is made: answers_the_ask reads the ask lexically, the judge tools refuse, and each result says so.`, ); } return { ok: false, reason }; @@ -239,6 +274,15 @@ export class JudgeBudget { this.ledger.settle(ticket.tenantId, ticket.day, -ticket.reservedMicroUsd, false); } + /** The same three steps for one tenant, as a judge tool or the citation verifier takes them around a call it makes itself. */ + gate(tenantId: TenantId): JudgeSpendGate { + return (worstCaseUsd) => { + const r = this.reserve(tenantId, worstCaseUsd); + if (!r.ok) return r; + return { ok: true, settle: (actualUsd) => this.settle(r.ticket, actualUsd), release: () => this.release(r.ticket) }; + }; + } + today(tenantId: TenantId): BudgetToday { const now = this.now(); const day = utcDay(now); @@ -274,3 +318,27 @@ export interface JudgeRequest { export function newJudgeRequest(): JudgeRequest { return { calls: 0, withheld: 0 }; } + +export interface JudgeBudgetEnvOptions { + /** Where the spend is kept: the server and the CLI pass the database's ledger, so it survives a restart. */ + ledger?: JudgeSpendLedger; + /** Where the one line goes when a tenant's budget first refuses a call on a day (default: stderr). */ + log?: (line: string) => void; + now?: () => Date; +} + +/** + * The one daily budget this process's environment sets, for every judge + * call it makes. `notes` are the settings that fell back or were read under + * their old name, one sentence each, for the process to say at startup. + */ +export function judgeBudgetFromEnv(options: JudgeBudgetEnvOptions = {}): { budget: JudgeBudget; notes: string[] } { + const daily = dailyBudgetUsd(); + const budget = new JudgeBudget({ + dailyUsd: daily.value, + ledger: options.ledger ?? memoryJudgeSpendLedger(), + log: options.log ?? ((line) => process.stderr.write(`${line}\n`)), + ...(options.now ? { now: options.now } : {}), + }); + return { budget, notes: daily.note ? [daily.note] : [] }; +} diff --git a/src/eval/llm-judge/relevance-judge.ts b/src/eval/llm-judge/relevance-judge.ts index cdbe88f0..c2d080d7 100644 --- a/src/eval/llm-judge/relevance-judge.ts +++ b/src/eval/llm-judge/relevance-judge.ts @@ -25,8 +25,8 @@ * * The user owns the key and the bill, so three things hold on every call * (budget.ts and redact.ts carry the reasoning): - * - a daily budget per tenant, kept in the database, that a call's worst - * case must fit before it is made; + * - a daily budget per tenant, kept in the database and shared with the + * judge tools, that a call's worst case must fit before it is made; * - a cap on judge calls per request, so one batch cannot spend the day; * - what leaves the machine is the ask and the answer with every span * no_pii flags replaced by a marker, unless the deployment opts out. @@ -45,12 +45,13 @@ import { JudgeBudget, MAX_CALLS_PER_REQUEST_VAR, dailyBudgetUsd, + judgeBudgetFromEnv, maxCallsPerRequest, memoryJudgeSpendLedger, newJudgeRequest, type BudgetToday, + type JudgeBudgetEnvOptions, type JudgeRequest, - type JudgeSpendLedger, type Setting, } from './budget.js'; import { redactForJudge } from './redact.js'; @@ -114,7 +115,7 @@ export interface RelevanceJudgeOptions { /** The judge call itself; tests and the proof runner pass the real evaluator or a stand-in. */ evaluate?: (params: LLMJudgeEvaluateParams) => Promise; /** - * The daily budget. Omitted: IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD (or its + * The daily budget. Omitted: IRIS_LLM_JUDGE_DAILY_BUDGET_USD (or its * default) over a ledger in this process's memory; an embedder with a * database passes one built on it, as the server and the CLI do. */ @@ -242,12 +243,13 @@ export function createRelevanceJudge(options: RelevanceJudgeOptions): RelevanceJ * IRIS_RELEVANCE_JUDGE_MODEL is unset. Literal reads on purpose: the docs * contract greps `process.env.IRIS_*` to learn what the server reads. */ -export interface RelevanceJudgeEnvOptions { - /** Where the daily spend is kept: the server and the CLI pass the database's ledger, so it survives a restart. */ - ledger?: JudgeSpendLedger; - /** Where the one line goes when a tenant's budget first refuses a call on a day (default: stderr). */ - log?: (line: string) => void; - now?: () => Date; +export interface RelevanceJudgeEnvOptions extends JudgeBudgetEnvOptions { + /** + * The daily budget every judge call in this process draws on: the server + * passes the one its judge tools use, so the three share one balance. + * Omitted, one is built from the environment and the other options. + */ + budget?: JudgeBudget; } export function relevanceJudgeFromEnv(options: RelevanceJudgeEnvOptions = {}): RelevanceJudge | null { @@ -255,22 +257,18 @@ export function relevanceJudgeFromEnv(options: RelevanceJudgeEnvOptions = {}): R if (!model) return null; const provider = findPricing(model)?.provider; const apiKey = provider === 'anthropic' ? process.env.IRIS_ANTHROPIC_API_KEY : provider === 'openai' ? process.env.IRIS_OPENAI_API_KEY : undefined; - const daily = dailyBudgetUsd(); const calls = maxCallsPerRequest(); const redaction = redactionSetting(); - const budget = new JudgeBudget({ - dailyUsd: daily.value, - ledger: options.ledger ?? memoryJudgeSpendLedger(), - log: options.log ?? ((line) => process.stderr.write(`${line}\n`)), - ...(options.now ? { now: options.now } : {}), - }); + // A budget passed in is the process's own, and whoever built it says its notes. + const built = options.budget ? null : judgeBudgetFromEnv(options); + const budget = options.budget ?? built!.budget; return createRelevanceJudge({ model, ...(apiKey ? { apiKey } : {}), budget, maxCallsPerRequest: calls.value, redact: redaction.value === 'on', - notes: [daily.note, calls.note, redaction.note].filter((n): n is string => n !== undefined), + notes: [...(built?.notes ?? []), calls.note, redaction.note].filter((n): n is string => n !== undefined), }); } diff --git a/src/index.ts b/src/index.ts index 715571fc..4f680a5c 100644 --- a/src/index.ts +++ b/src/index.ts @@ -274,8 +274,10 @@ Environment variables (CLI flags take precedence): IRIS_RELEVANCE_JUDGE_MODEL A priced model id: answers_the_ask then asks this judge on every evaluation that carries input, and gates on its verdict (off by default). Sends that input and output to the model's provider on your key - IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD What the relevance judge may spend per UTC day, per tenant, kept in the - database (default: 1). Past it, answers_the_ask reads the ask lexically + IRIS_LLM_JUDGE_DAILY_BUDGET_USD What every judge call together may spend per UTC day, per tenant: the + relevance judge, evaluate_with_llm_judge and verify_citations, kept in the + database (default: 1). Past it no judge call is made + IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD Its name before 0.21.0, read when IRIS_LLM_JUDGE_DAILY_BUDGET_USD is unset IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST Relevance judge calls one request may make (default: 20) IRIS_RELEVANCE_JUDGE_REDACT on (default): PII and credentials no_pii flags are replaced before the input and output are sent to the judge; off sends them as they are diff --git a/src/judge-enablement.json b/src/judge-enablement.json index 261f31b3..05cb947e 100644 --- a/src/judge-enablement.json +++ b/src/judge-enablement.json @@ -5,7 +5,7 @@ "Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the \"env\" block of the iris-eval entry in your MCP config — \"iris-eval\": { \"command\": \"npx\", \"args\": [\"-y\", \"@iris-eval/mcp-server\"], \"env\": { \"IRIS_ANTHROPIC_API_KEY\": \"sk-ant-...\" } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval.", "Restart the MCP session. A running process never sees a variable set after it started.", "Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard.", - "Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it.", - "Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why." + "Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it.", + "Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why." ] } diff --git a/src/server.ts b/src/server.ts index 9b72a2ec..108ff047 100644 --- a/src/server.ts +++ b/src/server.ts @@ -13,6 +13,7 @@ import { buildInstructions } from './instructions.js'; import { buildCapabilities, type Capabilities } from './capabilities.js'; import { judgeState } from './judge-enablement.js'; import { relevanceJudgeFromEnv, relevanceJudgeStartupWarnings, relevanceJudgeState } from './eval/llm-judge/relevance-judge.js'; +import { judgeBudgetFromEnv } from './eval/llm-judge/budget.js'; import type { StoreGate } from './storage/ready.js'; import { errorResult } from './tools/respond.js'; import { toIrisError } from './tools/errors.js'; @@ -36,7 +37,7 @@ export interface IrisServer { export interface IrisServerOptions { /** `demo` when the server runs against the disposable demo database. */ mode?: 'real' | 'demo'; - /** Where a warning line goes: the relevance judge's budget says here when it first refuses a call on a day. */ + /** Where a warning line goes: the daily judge budget says here when it first refuses a call on a day. */ warn?: (line: string) => void; /** Hold tool calls and resource reads until the store serves (storage/ready.ts); none when it serves from the start. */ gate?: StoreGate; @@ -87,10 +88,16 @@ export function createIrisServer( * its model (IRIS_RELEVANCE_JUDGE_MODEL). A key alone never installs it: * the key enables evaluate_with_llm_judge, which a caller invokes and pays * for per call, and must not start billing every evaluation on upgrade. - * Its daily budget is kept in this database, so a restart does not reset it. + * + * Every judge call on the user's key, its and the judge tools', draws on + * one daily budget kept in this database, so a restart does not reset it + * and an agent calling a judge tool in a loop stops where the operator said. */ const warn = options?.warn ?? ((line: string) => process.stderr.write(`${line}\n`)); - evalEngine.setRelevanceJudge(relevanceJudgeFromEnv({ ledger: storage.judgeSpendLedger(), log: warn })); + const judgeBudget = judgeBudgetFromEnv({ ledger: storage.judgeSpendLedger(), log: warn }); + for (const note of judgeBudget.notes) warn(`LLM judge: ${note}.`); + evalEngine.setJudgeBudget(judgeBudget.budget); + evalEngine.setRelevanceJudge(relevanceJudgeFromEnv({ budget: judgeBudget.budget })); // A judge that is configured and cannot run fails open; say so once, at startup, where the operator is looking. for (const line of relevanceJudgeStartupWarnings(evalEngine.relevanceJudgeInForce())) warn(line); // Caller can inject a shared rule store (e.g. index.ts passes the diff --git a/src/tools/errors.ts b/src/tools/errors.ts index 61da951d..6104e300 100644 --- a/src/tools/errors.ts +++ b/src/tools/errors.ts @@ -24,6 +24,7 @@ import { ZodError } from 'zod'; import { LLMJudgeError } from '../eval/llm-judge/client.js'; import { CAPABILITIES_RESOURCE_URI } from '../resources/uris.js'; import { JUDGE_COST_CAP_VAR } from '../judge-enablement.js'; +import { DAILY_BUDGET_VAR } from '../eval/llm-judge/budget.js'; export const ERROR_CODE_CATALOGUE = [ 'IRIS_INVALID_ARGUMENT', @@ -128,6 +129,18 @@ export function toIrisError(err: unknown): IrisError { }); } + if (name === 'DailyBudgetError') { + return irisError('IRIS_BUDGET_EXCEEDED', message, { + field: DAILY_BUDGET_VAR, + retryable: true, + recovery: [ + `Every judge call Iris makes on this key draws on one daily budget, the operator's ${DAILY_BUDGET_VAR}. Retry after it resets (the time is in the message), or the operator raises it.`, + 'Nothing was spent.', + ], + see: CAPABILITIES, + }); + } + if (name === 'CostCapError') { const e = err as { estimatedUsd?: number; capUsd?: number }; return irisError('IRIS_BUDGET_EXCEEDED', message, { diff --git a/src/tools/evaluate-with-llm-judge.ts b/src/tools/evaluate-with-llm-judge.ts index 0a67d84e..ade7d7b3 100644 --- a/src/tools/evaluate-with-llm-judge.ts +++ b/src/tools/evaluate-with-llm-judge.ts @@ -2,7 +2,7 @@ import { z } from 'zod'; import type { McpServer } from '@modelcontextprotocol/sdk/server/mcp.js'; import type { IStorageAdapter } from '../types/query.js'; import { LOCAL_TENANT } from '../types/tenant.js'; -import { evaluateWithLLMJudge } from '../eval/llm-judge/evaluator.js'; +import { CostCapError, evaluateWithLLMJudge, worstCaseJudgeCostUsd } from '../eval/llm-judge/evaluator.js'; import { judgeEvalResult } from '../eval/llm-judge/persisted.js'; import type { EvalEngine } from '../eval/engine.js'; import { verdictSchema } from '../eval/response-schema.js'; @@ -11,7 +11,7 @@ import type { LLMProvider } from '../eval/llm-judge/client.js'; import type { TemplateName } from '../eval/llm-judge/templates/index.js'; import { generateEvalId } from '../utils/ids.js'; import { JUDGE_COST_CAP_VAR, JUDGE_DEFAULT_COST_CAP_USD, JUDGE_KEY_VARS, judgeRecovery } from '../judge-enablement.js'; -import { asRecord, asWarning, judgeCostCeiling } from './operator-ceilings.js'; +import { DailyBudgetError, asRecord, asWarning, judgeBudgetFor, judgeCostCeiling } from './operator-ceilings.js'; import { strictInput } from './strict-input.js'; import { getTraceOrThrow, insertLinkedEvalResult } from './trace-link.js'; import { besideNote } from '../eval/of-record.js'; @@ -121,6 +121,7 @@ export function registerEvaluateWithLLMJudgeTool( storage: IStorageAdapter, engine: EvalEngine, ): void { + const judgeBudget = judgeBudgetFor(engine, storage); server.registerTool( 'evaluate_with_llm_judge', { @@ -182,7 +183,7 @@ export function registerEvaluateWithLLMJudgeTool( ...(agentModel !== null && sameFamily(args.model, agentModel) ? [sameFamilyWarning(args.model, agentModel)] : []), ]; - const result = await evaluateWithLLMJudge({ + const params = { output: args.output, template: args.template as TemplateName, model: args.model, @@ -195,7 +196,28 @@ export function registerEvaluateWithLLMJudgeTool( maxOutputTokens: args.max_output_tokens, temperature: args.temperature, timeoutMs: args.timeout_ms, - }); + }; + /* + * The daily budget every judge call on this key draws on: the call's + * worst case, priced as the per-call cap prices it, must fit in what is + * left today, or nothing is spent. A call over the per-call cap is not + * reserved: the evaluator refuses it before any spend. + */ + const worst = worstCaseJudgeCostUsd(params); + const hold = worst !== null && worst <= maxCostUsd ? judgeBudget().gate(LOCAL_TENANT)(worst) : null; + if (hold !== null && !hold.ok) throw new DailyBudgetError(hold.reason); + let result; + try { + result = await evaluateWithLLMJudge(params); + } catch (err) { + // Refused before the call, nothing was spent; any other failure may have been billed, so its worst case stays counted. + if (hold?.ok) { + if (err instanceof CostCapError) hold.release(); + else hold.settle(null); + } + throw err; + } + if (hold?.ok) hold.settle(result.costUsd); // Persist to eval_results so the dashboard can surface it. // eval_type='custom' — LLM judge scores span all 4 heuristic diff --git a/src/tools/guide.ts b/src/tools/guide.ts index ec1712c0..6b0b6d7c 100644 --- a/src/tools/guide.ts +++ b/src/tools/guide.ts @@ -14,6 +14,7 @@ import { RULE_ALPHA } from '../eval/compare.js'; import { INJECTION_SCOPE_SENTENCE } from '../eval/rules/safety.js'; import { JUDGE_COST_CAP_VAR, JUDGE_DEFAULT_COST_CAP_USD, JUDGE_KEY_VARS } from '../judge-enablement.js'; +import { DAILY_BUDGET_VAR } from '../eval/llm-judge/budget.js'; import { OUTPUT_SCHEMAS, TOOL_NAMES, type ToolName } from './index.js'; export interface ToolGuide { @@ -224,6 +225,7 @@ const TOOL_GUIDE: Record = { `Calls Anthropic or OpenAI directly with the key in this process's environment (${JUDGE_KEY_VARS.anthropic} or ${JUDGE_KEY_VARS.openai}); Iris never proxies. ` + 'template picks the question: accuracy, helpfulness, safety, correctness (needs expected), faithfulness (needs source_material), task_completed (pass the trajectory as source_material when you have it) or relevance (needs input: does the output address this request, not another); input improves helpfulness and safety. model is required; provider is inferred from it. ' + `The worst-case spend — both attempts, full max_output_tokens — is computed BEFORE the call and refused if it exceeds max_cost_usd (default ${JUDGE_COST_CAP_VAR} or ${JUDGE_DEFAULT_COST_CAP_USD}); max_cost_usd can lower the operator's cap, never raise it, and a larger value is lowered with an IRIS_ARGUMENT_NARROWED warning. ` + + `The same worst case must also fit in what is left of the day's ${DAILY_BUDGET_VAR}, which every judge call on the key shares. ` + 'temperature defaults to 0; a rate-limited call is retried once. One evaluation row is stored with the provider response id, tokens, cost and latency. With trace_id it is kept beside that trace (reference_trace_id) and listed with it; a judgment is never the trace\'s verdict and cannot replace it. ' + 'A judge from the same model family as the agent (agent_model, or the linked trace) is warned about, never refused. ' + "The judge's own accuracy is measurable on a key you supply and is not yet published (see iris://proof).", @@ -231,7 +233,7 @@ const TOOL_GUIDE: Record = { 'For length, keyword, PII, injection or cost checks: evaluate_output is free and deterministic. Without a key: the call returns IRIS_JUDGE_NOT_ENABLED with the enable steps — do not search for them. On very large outputs without raising max_cost_usd: the pre-check refuses.', errors: 'IRIS_INVALID_ARGUMENT for relevance without input, before any spend. IRIS_JUDGE_NOT_ENABLED (no key for the provider reached this process; recovery carries the steps). IRIS_JUDGE_UNKNOWN_MODEL (valid lists the models). IRIS_UNKNOWN_TRACE, checked before any spend. ' + - 'IRIS_BUDGET_EXCEEDED (nothing spent; the message carries both numbers). IRIS_PROVIDER_ERROR with kind auth, rate_limit, bad_request, server_error, timeout or malformed_response, and retryable set.', + `IRIS_BUDGET_EXCEEDED (nothing spent): over max_cost_usd, the message carries both numbers; over the daily budget, field is ${DAILY_BUDGET_VAR}, retryable is true and the message says when it resets. IRIS_PROVIDER_ERROR with kind auth, rate_limit, bad_request, server_error, timeout or malformed_response, and retryable set.`, parameters: { template: 'Judge dimension: accuracy (factual correctness), helpfulness (does it address the ask), safety (harm potential), correctness (vs reference answer — requires `expected`), faithfulness (RAG grounding — requires `source_material`), task_completed (did the task actually complete — pass the trajectory as `source_material` when you have it), relevance (does it address THIS request rather than another — requires `input`; the judge answers_the_ask gates on when IRIS_RELEVANCE_JUDGE_MODEL is set).', @@ -242,7 +244,8 @@ const TOOL_GUIDE: Record = { verify_citations: { does: 'Three phases. Extraction, no network: [N] references, (Author, Year), bare URLs and DOIs. Fetch of URL and DOI citations only when the operator set IRIS_CITATION_ALLOW_FETCH=1 (allow_fetch: false skips it for a call; true cannot turn it on), through a scheme allowlist, private and cloud-metadata address blocking, an optional hostname allowlist (domain_allowlist, merged with IRIS_CITATION_DOMAINS), a per-source timeout and byte cap, and at most three re-checked redirects. ' + - 'Then one judge call per resolved citation on your own key, reading the first part of each source, capped in total by max_cost_usd_total, which can lower the operator\'s IRIS_CITATION_MAX_COST_USD_TOTAL and never raise it. An argument that asks for more than the operator allows is narrowed, with an IRIS_ARGUMENT_NARROWED warning. Up to max_citations are verified; extras are skipped, not errored. ' + + 'Then one judge call per resolved citation on your own key, reading the first part of each source, capped in total by max_cost_usd_total, which can lower the operator\'s IRIS_CITATION_MAX_COST_USD_TOTAL and never raise it. An argument that asks for more than the operator allows is narrowed, with an IRIS_ARGUMENT_NARROWED warning. ' + + `Each judge call's worst case must also fit in what is left of the day's ${DAILY_BUDGET_VAR}, which every judge call on the key shares; past it the citation gets judge_error daily_budget_reached and later ones are not attempted. Up to max_citations are verified; extras are skipped, not errored. ` + 'overall_score is supported / judged and null when nothing was judged. Per-citation failures are reported on the citation, never scored as unsupported: resolve_error when the source was not resolved (fetch disabled, bad scheme, blocked address, fetch timeout, bad status), judge_error when it was and the judge gave no verdict (cost cap, provider error, unreadable reply). One evaluation row is stored.', whenNot: "When the output has no citations: the score is null, and evaluate_output's hallucination signals are the cheap check. " + diff --git a/src/tools/operator-ceilings.ts b/src/tools/operator-ceilings.ts index 154649a5..efec580d 100644 --- a/src/tools/operator-ceilings.ts +++ b/src/tools/operator-ceilings.ts @@ -24,6 +24,9 @@ */ import { JUDGE_COST_CAP_VAR, judgeCostCapUsd } from '../judge-enablement.js'; import type { NarrowedArgument } from '../types/eval.js'; +import type { EvalEngine } from '../eval/engine.js'; +import type { IStorageAdapter } from '../types/query.js'; +import { judgeBudgetFromEnv, type JudgeBudget } from '../eval/llm-judge/budget.js'; export const CITATION_ALLOW_FETCH_VAR = 'IRIS_CITATION_ALLOW_FETCH'; export const CITATION_DOMAINS_VAR = 'IRIS_CITATION_DOMAINS'; @@ -152,3 +155,22 @@ export function judgeCostCeiling(asked: number | undefined): { value: number; wa export function citationCostCeiling(asked: number | undefined): { value: number; warning?: ArgumentNarrowed } { return costCeiling('max_cost_usd_total', asked, citationTotalCostCapUsd(), CITATION_TOTAL_COST_VAR); } + +/** + * The daily budget a judge tool draws on (llm-judge/budget.ts): the one the + * server set on the engine and shares with the relevance judge. A tool + * registered without one draws on the environment's over this store's + * ledger, which holds the same balance. + */ +export function judgeBudgetFor(engine: EvalEngine, storage: IStorageAdapter): () => JudgeBudget { + let own: JudgeBudget | null = null; + return () => engine.judgeBudgetInForce() ?? (own ??= judgeBudgetFromEnv({ ledger: storage.judgeSpendLedger() }).budget); +} + +/** A judge call the daily budget refused, before any spend: IRIS_BUDGET_EXCEEDED, retryable once the budget resets. */ +export class DailyBudgetError extends Error { + constructor(reason: string) { + super(`Not judged: ${reason}.`); + this.name = 'DailyBudgetError'; + } +} diff --git a/src/tools/verify-citations.ts b/src/tools/verify-citations.ts index d46b3f49..9b94aead 100644 --- a/src/tools/verify-citations.ts +++ b/src/tools/verify-citations.ts @@ -7,6 +7,7 @@ import { verifyCitations } from '../eval/citation-verify/verifier.js'; import type { LLMProvider } from '../eval/llm-judge/client.js'; import { generateEvalId } from '../utils/ids.js'; import { JUDGE_KEY_VARS } from '../judge-enablement.js'; +import { DAILY_BUDGET_VAR } from '../eval/llm-judge/budget.js'; import { strictInput } from './strict-input.js'; import { assertTraceExists, insertLinkedEvalResult } from './trace-link.js'; import { besideNote } from '../eval/of-record.js'; @@ -18,7 +19,7 @@ import { describeTool, ERROR_ENVELOPE_SENTENCE } from './describe.js'; import { advertisedOutput } from './advertise.js'; import { irisError } from './errors.js'; import { evaluationLinks, guarded, respond } from './respond.js'; -import { asRecord, asWarning, citationCostCeiling, domainCeiling, fetchCeiling } from './operator-ceilings.js'; +import { asRecord, asWarning, citationCostCeiling, domainCeiling, fetchCeiling, judgeBudgetFor } from './operator-ceilings.js'; const inputSchema = { output: z.string().min(1).describe('The agent output containing citations to verify'), @@ -76,11 +77,12 @@ export function assertJudgeRan(result: { `verify_citations could not judge any of the ${result.totalResolved} resolved citation(s): the judge failed on every one (${kinds}). ` + `Nothing was verified and nothing was stored, so there is no verdict. First error: ${first}`, { - retryable: /timeout|rate_limit|server_error/.test(kinds), + retryable: /timeout|rate_limit|server_error|daily_budget_reached/.test(kinds), recovery: [ 'Check the key and the model: a refused key or an unknown model fails every citation the same way.', 'Retry when the kind is a timeout, a rate limit or a provider server error.', 'Raise max_cost_usd_total when the kind is cost_cap_reached.', + `When the kind is daily_budget_reached, every judge call on this key has used the operator's ${DAILY_BUDGET_VAR} for today: retry after it resets, or the operator raises it.`, ], }, ); @@ -144,6 +146,7 @@ export const verifyCitationsOutputSchema = z.looseObject({ }); export function registerVerifyCitationsTool(server: McpServer, storage: IStorageAdapter, engine: EvalEngine): void { + const judgeBudget = judgeBudgetFor(engine, storage); server.registerTool( 'verify_citations', { @@ -197,6 +200,7 @@ export function registerVerifyCitationsTool(server: McpServer, storage: IStorage maxCitations: args.max_citations, perSourceTimeoutMs: args.per_source_timeout_ms, perSourceMaxBytes: args.per_source_max_bytes, + spend: judgeBudget().gate(LOCAL_TENANT), }); assertJudgeRan(result); diff --git a/tests/integration/relevance-judge-limits.test.ts b/tests/integration/relevance-judge-limits.test.ts index 2b1882ab..7f6dd8b0 100644 --- a/tests/integration/relevance-judge-limits.test.ts +++ b/tests/integration/relevance-judge-limits.test.ts @@ -49,6 +49,7 @@ const ENV = [ 'IRIS_RELEVANCE_JUDGE_MODEL', 'IRIS_ANTHROPIC_API_KEY', 'IRIS_OPENAI_API_KEY', + 'IRIS_LLM_JUDGE_DAILY_BUDGET_USD', 'IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD', 'IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST', 'IRIS_RELEVANCE_JUDGE_REDACT', @@ -155,7 +156,7 @@ const WORST = worstCaseJudgeCostUsd({ template: 'relevance', model: MODEL, input describe('the relevance judge daily budget', () => { it('admits calls while the worst case fits, then withholds the judge, falls back to the lexical reading, logs once and shows it on health', async () => { // Room for exactly two calls: the second needs one actual plus one worst case, the third two actuals plus one. - env({ IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD: String(2 * ACTUAL + WORST - 0.000002) }); + env({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: String(2 * ACTUAL + WORST - 0.000002) }); const storage = await store(); const warn = vi.fn(); const { server, client } = await mcp(storage, warn); @@ -169,8 +170,8 @@ describe('the relevance judge daily budget', () => { const third = await evaluate(client, ASK); expect(providerCalls).toHaveLength(2); expect(third.judge).toMatchObject({ withheld: 'daily_budget', costUsd: 0 }); - expect(third.judge?.error).toMatch(/daily budget has 0\.\d{4} of [\d.]+ USD left today \(UTC\)/); - expect(third.judge?.error).toContain('IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD'); + expect(third.judge?.error).toMatch(/daily judge budget has 0\.\d{4} of [\d.]+ USD left today \(UTC\)/); + expect(third.judge?.error).toContain('IRIS_LLM_JUDGE_DAILY_BUDGET_USD'); // The lexical reading decided, and it advises rather than gates. expect(third.kind).toBe('policy'); expect(third.role).toBe('advisory'); @@ -191,7 +192,7 @@ describe('the relevance judge daily budget', () => { }); it('is kept in the database: a new server on the same file starts from what was spent, not from zero', async () => { - env({ IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD: String(2 * ACTUAL + WORST - 0.000002) }); + env({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: String(2 * ACTUAL + WORST - 0.000002) }); const file = join(dir, 'restart.db'); { const storage = await store(file); @@ -235,7 +236,7 @@ describe('the relevance judge daily budget', () => { }, 60_000); it('0 stops every call before any spend', async () => { - env({ IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD: '0' }); + env({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: '0' }); const { client } = await mcp(await store()); expect((await evaluate(client, ASK)).judge?.withheld).toBe('daily_budget'); expect(providerCalls).toHaveLength(0); diff --git a/tests/mcpb-manifest.test.ts b/tests/mcpb-manifest.test.ts index 51b52500..7545fa7e 100644 --- a/tests/mcpb-manifest.test.ts +++ b/tests/mcpb-manifest.test.ts @@ -145,8 +145,8 @@ describe('mcpb/manifest.json', () => { // The relevance judge: off until a model is named, and inside the same daily budget the server defaults to. expect(manifest.user_config.relevance_judge_model).toMatchObject({ type: 'string', required: false, default: '' }); expect(manifest.server.mcp_config.env.IRIS_RELEVANCE_JUDGE_MODEL).toBe('${user_config.relevance_judge_model}'); - expect(manifest.user_config.relevance_judge_daily_budget_usd).toMatchObject({ type: 'number', required: false, default: DEFAULT_DAILY_BUDGET_USD }); - expect(manifest.server.mcp_config.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD).toBe('${user_config.relevance_judge_daily_budget_usd}'); + expect(manifest.user_config.judge_daily_budget_usd).toMatchObject({ type: 'number', required: false, default: DEFAULT_DAILY_BUDGET_USD }); + expect(manifest.server.mcp_config.env.IRIS_LLM_JUDGE_DAILY_BUDGET_USD).toBe('${user_config.judge_daily_budget_usd}'); expect(manifest.server.mcp_config.env.IRIS_DASHBOARD).toBe('${user_config.dashboard}'); for (const url of manifest.privacy_policies) expect(url.startsWith('https://'), url).toBe(true); }); diff --git a/tests/unit/eval/llm-judge/judge-budget.test.ts b/tests/unit/eval/llm-judge/judge-budget.test.ts index df54f314..f9edcfbe 100644 --- a/tests/unit/eval/llm-judge/judge-budget.test.ts +++ b/tests/unit/eval/llm-judge/judge-budget.test.ts @@ -1,5 +1,5 @@ /* - * The relevance judge's daily budget and its ledger, below the servers: + * The daily judge budget and its ledger, below the servers: * the ceiling is never passed, a call whose cost is unknown stays counted * at its worst case, a refusal before any spend gives the reservation back, * the day turns over at 00:00 UTC, tenants are kept apart, and two @@ -14,6 +14,7 @@ import { DEFAULT_MAX_CALLS_PER_REQUEST, JudgeBudget, dailyBudgetUsd, + judgeBudgetFromEnv, maxCallsPerRequest, memoryJudgeSpendLedger, nextUtcMidnight, @@ -98,7 +99,7 @@ describe('JudgeBudget', () => { }); describe('the settings', () => { - const vars = ['IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD', 'IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST'] as const; + const vars = ['IRIS_LLM_JUDGE_DAILY_BUDGET_USD', 'IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD', 'IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST'] as const; const saved = Object.fromEntries(vars.map((k) => [k, process.env[k]])); afterEach(() => { for (const k of vars) { @@ -114,18 +115,73 @@ describe('the settings', () => { expect(DEFAULT_DAILY_BUDGET_USD).toBe(1); expect(DEFAULT_MAX_CALLS_PER_REQUEST).toBe(20); - process.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD = '2.5'; + process.env.IRIS_LLM_JUDGE_DAILY_BUDGET_USD = '2.5'; process.env.IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST = '0'; expect(dailyBudgetUsd()).toEqual({ value: 2.5, source: 'env' }); expect(maxCallsPerRequest()).toEqual({ value: 0, source: 'env' }); - process.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD = '5$'; + process.env.IRIS_LLM_JUDGE_DAILY_BUDGET_USD = '5$'; process.env.IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST = '2.5'; expect(dailyBudgetUsd()).toMatchObject({ value: 1, source: 'default', note: expect.stringMatching(/"5\$" is not a number/) }); expect(maxCallsPerRequest()).toMatchObject({ value: 20, source: 'default', note: expect.stringMatching(/"2\.5" is not a whole number/) }); }); }); +describe('the old name of the daily budget', () => { + const vars = ['IRIS_LLM_JUDGE_DAILY_BUDGET_USD', 'IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD'] as const; + const saved = Object.fromEntries(vars.map((k) => [k, process.env[k]])); + afterEach(() => { + for (const k of vars) { + if (saved[k] === undefined) delete process.env[k]; + else process.env[k] = saved[k]; + } + }); + + it('is read when the new name is unset, and says it now limits every judge call', () => { + delete process.env.IRIS_LLM_JUDGE_DAILY_BUDGET_USD; + process.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD = '3'; + const s = dailyBudgetUsd(); + expect(s).toMatchObject({ value: 3, source: 'env' }); + expect(s.note).toContain('IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD is the old name of IRIS_LLM_JUDGE_DAILY_BUDGET_USD'); + expect(judgeBudgetFromEnv().notes).toEqual([s.note]); + }); + + it('gives way to the new name, silently', () => { + process.env.IRIS_LLM_JUDGE_DAILY_BUDGET_USD = '0.5'; + process.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD = '3'; + expect(dailyBudgetUsd()).toEqual({ value: 0.5, source: 'env' }); + expect(judgeBudgetFromEnv().notes).toEqual([]); + }); + + it('a bad value under the old name is named by that name', () => { + delete process.env.IRIS_LLM_JUDGE_DAILY_BUDGET_USD; + process.env.IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD = 'lots'; + expect(dailyBudgetUsd()).toMatchObject({ value: 1, source: 'default', note: expect.stringMatching(/^IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD="lots"/) }); + }); +}); + +describe('a gate for a caller that makes the call itself', () => { + it('holds the worst case, settles to the cost, releases a call never made, and refuses past the limit', () => { + const budget = new JudgeBudget({ dailyUsd: 0.01, ledger: memoryJudgeSpendLedger(), now: () => new Date('2026-10-05T12:00:00Z') }); + const gate = budget.gate(LOCAL_TENANT); + const a = gate(0.006); + expect(a.ok).toBe(true); + if (a.ok) a.settle(0.001); + expect(budget.today(LOCAL_TENANT)).toMatchObject({ spentUsd: 0.001, calls: 1 }); + const b = gate(0.006); + expect(b.ok).toBe(true); + if (b.ok) b.release(); + expect(budget.today(LOCAL_TENANT)).toMatchObject({ spentUsd: 0.001, calls: 1 }); + // A call that failed after the provider may have billed it keeps its worst case. + const c = gate(0.004); + if (c.ok) c.settle(null); + expect(budget.today(LOCAL_TENANT)).toMatchObject({ spentUsd: 0.005, calls: 2 }); + const d = gate(0.006); + expect(d.ok).toBe(false); + if (!d.ok) expect(d.reason).toMatch(/daily judge budget has 0\.0050 of 0\.01 USD left today \(UTC\).*IRIS_LLM_JUDGE_DAILY_BUDGET_USD raises it/); + }); +}); + describe('the SQLite ledger', () => { let dir: string; const adapters: SqliteAdapter[] = []; diff --git a/tests/unit/tools/judge-daily-budget.test.ts b/tests/unit/tools/judge-daily-budget.test.ts new file mode 100644 index 00000000..661d6976 --- /dev/null +++ b/tests/unit/tools/judge-daily-budget.test.ts @@ -0,0 +1,130 @@ +/* + * One daily budget for every judge call on the user's key. + * + * The relevance judge had a daily budget; evaluate_with_llm_judge and + * verify_citations had a per-call cap and nothing over the day, so an agent + * calling either in a loop could spend the key without end. These run both + * tools over an in-memory MCP transport against the budget the server sets + * from IRIS_LLM_JUDGE_DAILY_BUDGET_USD, and check that a refused call spends + * nothing, that both tools draw on one balance, and that the relevance judge + * draws on the same one. + * + * The provider client is mocked, so nothing is spent; the cited host is + * answered by a stubbed fetch, and DNS by a stub. + */ +import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'; +import { Client } from '@modelcontextprotocol/sdk/client/index.js'; +import { InMemoryTransport } from '@modelcontextprotocol/sdk/inMemory.js'; + +const callLLMJudge = vi.fn(); +vi.mock('../../../src/eval/llm-judge/client.js', () => ({ + callLLMJudge: (...args: unknown[]) => callLLMJudge(...args) as unknown, + estimateInputTokens: () => 100, + LLMJudgeError: class extends Error { + kind = 'server_error'; + retryable = false; + }, +})); + +const { SqliteAdapter } = await import('../../../src/storage/sqlite-adapter.js'); +const { createIrisServer } = await import('../../../src/server.js'); +const { defaultConfig } = await import('../../../src/config/defaults.js'); +const { LOCAL_TENANT } = await import('../../../src/types/tenant.js'); +const { __clearCitationCacheForTests, __setDnsLookupForTests } = await import('../../../src/eval/citation-verify/resolve.js'); + +type Result = { content?: unknown; isError?: boolean }; +type ErrorBody = { error: { code: string; message: string; retryable: boolean; field?: string } }; +const body = (r: Result) => + JSON.parse((r.content as Array<{ type: string; text: string }>).find((c) => c.type === 'text')!.text) as Record; + +const SCORE = JSON.stringify({ score: 0.9, rationale: 'because', dimensions: { a: 0.9 } }); +const SUPPORTED = '{"supported":true,"confidence":0.9,"rationale":"the page states it"}'; +const reply = (content: string) => ({ content, inputTokens: 10, outputTokens: 10, latencyMs: 1, rawProviderResponseId: 'resp-1' }); + +describe('every judge call on the key draws on one daily budget', () => { + let storage: InstanceType; + let server: ReturnType; + let client: Client; + const savedFetch = global.fetch; + + /** The server reads the budget from the environment when it is created, so each case sets it first. */ + async function start(env: Record): Promise { + for (const [k, v] of Object.entries(env)) vi.stubEnv(k, v); + storage = new SqliteAdapter(':memory:'); + await storage.initialize(); + server = createIrisServer(defaultConfig, storage, undefined, { warn: () => {} }); + const [c, s] = InMemoryTransport.createLinkedPair(); + await server.mcpServer.connect(s); + client = new Client({ name: 'judge-daily-budget', version: '0.1.0' }); + await client.connect(c); + } + + const judge = async () => + body((await client.callTool({ name: 'evaluate_with_llm_judge', arguments: { output: 'an answer', template: 'accuracy', model: 'claude-haiku-4-5' } })) as Result); + const verify = async () => + body((await client.callTool({ name: 'verify_citations', arguments: { output: 'Alpha is 42 (https://a.example/alpha).', model: 'claude-haiku-4-5' } })) as Result); + + beforeEach(() => { + callLLMJudge.mockReset(); + vi.stubEnv('IRIS_ANTHROPIC_API_KEY', 'sk-ant-dummy-key-for-tests-0123456789'); + vi.stubEnv('IRIS_CITATION_ALLOW_FETCH', '1'); + vi.stubEnv('IRIS_RELEVANCE_JUDGE_MODEL', ''); + vi.stubEnv('IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD', ''); + __setDnsLookupForTests(async () => [{ address: '93.184.216.34', family: 4 }]); + global.fetch = vi.fn(async () => new Response('Alpha is 42.', { status: 200, headers: { 'content-type': 'text/plain' } })) as unknown as typeof fetch; + }); + + afterEach(async () => { + await client.close(); + await storage.close(); + global.fetch = savedFetch; + __clearCitationCacheForTests(); + __setDnsLookupForTests(null); + vi.unstubAllEnvs(); + }); + + it('a spent budget refuses evaluate_with_llm_judge before any call: IRIS_BUDGET_EXCEEDED, retryable, nothing stored', async () => { + await start({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: '0' }); + const out = (await judge()) as unknown as ErrorBody; + expect(callLLMJudge).not.toHaveBeenCalled(); + expect(out.error).toMatchObject({ code: 'IRIS_BUDGET_EXCEEDED', retryable: true, field: 'IRIS_LLM_JUDGE_DAILY_BUDGET_USD' }); + expect(out.error.message).toMatch(/daily judge budget has 0\.0000 of 0 USD left today \(UTC\).*resets at /); + expect((await storage.queryEvalResults(LOCAL_TENANT, {})).total).toBe(0); + }); + + it('a spent budget refuses every citation judge call: daily_budget_reached, and the tool fails closed as retryable', async () => { + await start({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: '0' }); + const out = (await verify()) as unknown as ErrorBody; + expect(callLLMJudge).not.toHaveBeenCalled(); + expect(out.error).toMatchObject({ code: 'IRIS_JUDGE_FAILED', retryable: true }); + expect(out.error.message).toContain('daily_budget_reached'); + }); + + it('both tools draw on one balance, and today\'s spend is on iris://capabilities', async () => { + await start({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: '0.5' }); + callLLMJudge.mockResolvedValueOnce(reply(SCORE)).mockResolvedValueOnce(reply(SUPPORTED)); + expect((await judge()).passed).toBeDefined(); + expect((await verify()).total_judged).toBe(1); + expect(callLLMJudge).toHaveBeenCalledTimes(2); + + const today = server.capabilities().judge.dailyBudget; + expect(today).toMatchObject({ limitUsd: 0.5, calls: 2, refused: 0, exhausted: false }); + // Settled to what the calls cost, not held at their worst case. + expect(today!.spentUsd).toBeGreaterThan(0); + expect(today!.spentUsd).toBeLessThan(0.001); + }); + + it('the relevance judge draws on the same budget the tools do', async () => { + await start({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: '0.5', IRIS_RELEVANCE_JUDGE_MODEL: 'claude-haiku-4-5' }); + const relevance = server.evalEngine.relevanceJudgeInForce(); + expect(relevance).not.toBeNull(); + expect(relevance!.budget).toBe(server.evalEngine.judgeBudgetInForce()); + }); + + it('the old variable name still sets the budget when the new one is unset', async () => { + await start({ IRIS_LLM_JUDGE_DAILY_BUDGET_USD: '', IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD: '0' }); + const out = (await judge()) as unknown as ErrorBody; + expect(callLLMJudge).not.toHaveBeenCalled(); + expect(out.error.code).toBe('IRIS_BUDGET_EXCEEDED'); + }); +}); diff --git a/website/public/llms-full.txt b/website/public/llms-full.txt index f81c2b6f..79f5361b 100644 --- a/website/public/llms-full.txt +++ b/website/public/llms-full.txt @@ -111,11 +111,11 @@ Every snippet here pins the current release, so a config copied today keeps runn - When another call is better: To expire old data in bulk (retention.days; the sweep runs at boot and every retention.sweepIntervalHours). To delete evaluations: they are not deleted per row; retention and --purge cover them. To pause anything: traces are immutable, there is nothing to pause. - Errors: IRIS_STORAGE_ERROR when the delete cannot run. A malformed trace_id (not 32 lowercase hex) is refused before the handler runs. 11. `evaluate_with_llm_judge` — Score an output with an LLM judge on your own key: a 0..1 score, a rationale, the spend. - - What it does: Calls Anthropic or OpenAI directly with the key in this process's environment (IRIS_ANTHROPIC_API_KEY or IRIS_OPENAI_API_KEY); Iris never proxies. template picks the question: accuracy, helpfulness, safety, correctness (needs expected), faithfulness (needs source_material), task_completed (pass the trajectory as source_material when you have it) or relevance (needs input: does the output address this request, not another); input improves helpfulness and safety. model is required; provider is inferred from it. The worst-case spend — both attempts, full max_output_tokens — is computed BEFORE the call and refused if it exceeds max_cost_usd (default IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL or 0.25); max_cost_usd can lower the operator's cap, never raise it, and a larger value is lowered with an IRIS_ARGUMENT_NARROWED warning. temperature defaults to 0; a rate-limited call is retried once. One evaluation row is stored with the provider response id, tokens, cost and latency. With trace_id it is kept beside that trace (reference_trace_id) and listed with it; a judgment is never the trace's verdict and cannot replace it. A judge from the same model family as the agent (agent_model, or the linked trace) is warned about, never refused. The judge's own accuracy is measurable on a key you supply and is not yet published (see iris://proof). + - What it does: Calls Anthropic or OpenAI directly with the key in this process's environment (IRIS_ANTHROPIC_API_KEY or IRIS_OPENAI_API_KEY); Iris never proxies. template picks the question: accuracy, helpfulness, safety, correctness (needs expected), faithfulness (needs source_material), task_completed (pass the trajectory as source_material when you have it) or relevance (needs input: does the output address this request, not another); input improves helpfulness and safety. model is required; provider is inferred from it. The worst-case spend — both attempts, full max_output_tokens — is computed BEFORE the call and refused if it exceeds max_cost_usd (default IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL or 0.25); max_cost_usd can lower the operator's cap, never raise it, and a larger value is lowered with an IRIS_ARGUMENT_NARROWED warning. The same worst case must also fit in what is left of the day's IRIS_LLM_JUDGE_DAILY_BUDGET_USD, which every judge call on the key shares. temperature defaults to 0; a rate-limited call is retried once. One evaluation row is stored with the provider response id, tokens, cost and latency. With trace_id it is kept beside that trace (reference_trace_id) and listed with it; a judgment is never the trace's verdict and cannot replace it. A judge from the same model family as the agent (agent_model, or the linked trace) is warned about, never refused. The judge's own accuracy is measurable on a key you supply and is not yet published (see iris://proof). - When another call is better: For length, keyword, PII, injection or cost checks: evaluate_output is free and deterministic. Without a key: the call returns IRIS_JUDGE_NOT_ENABLED with the enable steps — do not search for them. On very large outputs without raising max_cost_usd: the pre-check refuses. - - Errors: IRIS_INVALID_ARGUMENT for relevance without input, before any spend. IRIS_JUDGE_NOT_ENABLED (no key for the provider reached this process; recovery carries the steps). IRIS_JUDGE_UNKNOWN_MODEL (valid lists the models). IRIS_UNKNOWN_TRACE, checked before any spend. IRIS_BUDGET_EXCEEDED (nothing spent; the message carries both numbers). IRIS_PROVIDER_ERROR with kind auth, rate_limit, bad_request, server_error, timeout or malformed_response, and retryable set. + - Errors: IRIS_INVALID_ARGUMENT for relevance without input, before any spend. IRIS_JUDGE_NOT_ENABLED (no key for the provider reached this process; recovery carries the steps). IRIS_JUDGE_UNKNOWN_MODEL (valid lists the models). IRIS_UNKNOWN_TRACE, checked before any spend. IRIS_BUDGET_EXCEEDED (nothing spent): over max_cost_usd, the message carries both numbers; over the daily budget, field is IRIS_LLM_JUDGE_DAILY_BUDGET_USD, retryable is true and the message says when it resets. IRIS_PROVIDER_ERROR with kind auth, rate_limit, bad_request, server_error, timeout or malformed_response, and retryable set. 12. `verify_citations` — Is each citation in an output supported by its source? Extracts, fetches (opt-in, SSRF-guarded) and judges on your key. - - What it does: Three phases. Extraction, no network: [N] references, (Author, Year), bare URLs and DOIs. Fetch of URL and DOI citations only when the operator set IRIS_CITATION_ALLOW_FETCH=1 (allow_fetch: false skips it for a call; true cannot turn it on), through a scheme allowlist, private and cloud-metadata address blocking, an optional hostname allowlist (domain_allowlist, merged with IRIS_CITATION_DOMAINS), a per-source timeout and byte cap, and at most three re-checked redirects. Then one judge call per resolved citation on your own key, reading the first part of each source, capped in total by max_cost_usd_total, which can lower the operator's IRIS_CITATION_MAX_COST_USD_TOTAL and never raise it. An argument that asks for more than the operator allows is narrowed, with an IRIS_ARGUMENT_NARROWED warning. Up to max_citations are verified; extras are skipped, not errored. overall_score is supported / judged and null when nothing was judged. Per-citation failures are reported on the citation, never scored as unsupported: resolve_error when the source was not resolved (fetch disabled, bad scheme, blocked address, fetch timeout, bad status), judge_error when it was and the judge gave no verdict (cost cap, provider error, unreadable reply). One evaluation row is stored. + - What it does: Three phases. Extraction, no network: [N] references, (Author, Year), bare URLs and DOIs. Fetch of URL and DOI citations only when the operator set IRIS_CITATION_ALLOW_FETCH=1 (allow_fetch: false skips it for a call; true cannot turn it on), through a scheme allowlist, private and cloud-metadata address blocking, an optional hostname allowlist (domain_allowlist, merged with IRIS_CITATION_DOMAINS), a per-source timeout and byte cap, and at most three re-checked redirects. Then one judge call per resolved citation on your own key, reading the first part of each source, capped in total by max_cost_usd_total, which can lower the operator's IRIS_CITATION_MAX_COST_USD_TOTAL and never raise it. An argument that asks for more than the operator allows is narrowed, with an IRIS_ARGUMENT_NARROWED warning. Each judge call's worst case must also fit in what is left of the day's IRIS_LLM_JUDGE_DAILY_BUDGET_USD, which every judge call on the key shares; past it the citation gets judge_error daily_budget_reached and later ones are not attempted. Up to max_citations are verified; extras are skipped, not errored. overall_score is supported / judged and null when nothing was judged. Per-citation failures are reported on the citation, never scored as unsupported: resolve_error when the source was not resolved (fetch disabled, bad scheme, blocked address, fetch timeout, bad status), judge_error when it was and the judge gave no verdict (cost cap, provider error, unreadable reply). One evaluation row is stored. - When another call is better: When the output has no citations: the score is null, and evaluate_output's hallucination signals are the cheap check. Without a key (IRIS_ANTHROPIC_API_KEY or IRIS_OPENAI_API_KEY): the call returns IRIS_JUDGE_NOT_ENABLED with the enable steps. With fetch enabled and an open allowlist on untrusted output: you are running a user-directed fetcher — set IRIS_CITATION_DOMAINS. - Errors: IRIS_JUDGE_NOT_ENABLED, IRIS_JUDGE_UNKNOWN_MODEL and IRIS_UNKNOWN_TRACE before any fetch or spend. IRIS_JUDGE_FAILED when citations resolved but the judge failed on every one — an error, not a passing verdict; nothing is stored.