Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 7 additions & 7 deletions .claims.json
Original file line number Diff line number Diff line change
Expand Up @@ -5141,8 +5141,8 @@
],
"version": 1
},
"generatedAt": "2026-10-05T15:25:24.087Z",
"generatedFromCommit": "42365584",
"generatedAt": "2026-10-05T17:06:12.498Z",
"generatedFromCommit": "547c3c39",
"generatorVersion": "1.0.0",
"llmJudgeTemplates": {
"count": 7,
Expand All @@ -5152,8 +5152,8 @@
"Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the \"env\" block of the iris-eval entry in your MCP config — \"iris-eval\": { \"command\": \"npx\", \"args\": [\"-y\", \"@iris-eval/mcp-server\"], \"env\": { \"IRIS_ANTHROPIC_API_KEY\": \"sk-ant-...\" } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval.",
"Restart the MCP session. A running process never sees a variable set after it started.",
"Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard.",
"Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it.",
"Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why."
"Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it.",
"Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why."
],
"title": "Enable the LLM judge (optional; the deterministic rules never need it)"
},
Expand Down Expand Up @@ -9497,16 +9497,16 @@
"passed": null,
"total": 51
},
"totalCombined": 5298,
"totalCombined": 5307,
"vitestDashboard": {
"failed": 0,
"passed": 443,
"total": 443
},
"vitestRoot": {
"failed": 0,
"passed": 4804,
"total": 4804
"passed": 4813,
"total": 4813
}
},
"version": {
Expand Down
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Security

- **BREAKING — A tool argument can no longer widen what the operator allows for fetching or spend.** Three arguments could. `verify_citations` with `allow_fetch: true` fetched the cited URLs on a server whose operator had not set `IRIS_CITATION_ALLOW_FETCH=1`; its `domain_allowlist` was added to the operator's `IRIS_CITATION_DOMAINS` instead of narrowing it; and `max_cost_usd` on `evaluate_with_llm_judge` and `max_cost_usd_total` on `verify_citations` replaced the operator's cap with any larger number. An agent's arguments can be steered by text it read, including text stored in Iris, so each is now a ceiling: an argument can narrow the operator's setting for one call and never widen it. When one asks for more, the operator's setting applies and the call goes on, the response carries a warning with code `IRIS_ARGUMENT_NARROWED` naming the argument and the setting that applied, and the stored evaluation keeps the same fact (`provenance.narrowed`) with a sentence to the operator in its `interpretations`, on every later read. Fetching now needs `IRIS_CITATION_ALLOW_FETCH=1` on the server; a caller that relied on `allow_fetch: true` gets no fetch and the warning. The per-call cap on `verify_citations` is a new setting, `IRIS_CITATION_MAX_COST_USD_TOTAL`, default 1 USD as before.
- **BREAKING — Every judge call on your key draws on one daily budget.** `evaluate_with_llm_judge` and `verify_citations` had a per-call cap and no daily limit, so an agent calling either in a loop, or steered into doing so, could spend the key without end; only the relevance judge had a daily budget. All three now share one, `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` (default 1 USD per UTC day, per tenant, kept in the database). A call is made only if its worst case fits in what is left. When it does not, nothing is spent: `evaluate_with_llm_judge` answers `IRIS_BUDGET_EXCEEDED` (retryable, with the time the budget resets), and `verify_citations` marks the citation `daily_budget_reached` and makes no further call. A deployment that spends more than 1 USD a day on the judge tools raises the budget. The variable was `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD`; that name is still read when the new one is unset, and the server says so at startup. In the Claude Desktop extension the setting is now "LLM judge daily budget", and a value set under the old one is not carried over. Today's spend is `judge.dailyBudget` on `iris://capabilities`.

## [0.20.0] - 2026-10-03

Expand Down
7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -354,8 +354,8 @@ Iris registers twelve tools that any MCP-compatible agent can invoke — trace a
2. Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the "env" block of the iris-eval entry in your MCP config — "iris-eval": { "command": "npx", "args": ["-y", "@iris-eval/mcp-server"], "env": { "IRIS_ANTHROPIC_API_KEY": "sk-ant-..." } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval.
3. Restart the MCP session. A running process never sees a variable set after it started.
4. Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard.
5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it.
6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why.
5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it.
6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why.

When `IRIS_OTEL_ENDPOINT` is configured, `log_trace` calls also emit a best-effort OTLP/HTTP JSON export to any OpenTelemetry collector (Jaeger, Grafana Tempo, Datadog OTLP, Honeycomb, etc). See [docs/otel-integration.md](https://github.com/iris-eval/mcp-server/blob/main/docs/otel-integration.md).

Expand Down Expand Up @@ -485,7 +485,8 @@ Every variable `--help` documents. CLI flags take precedence over environment va
| `IRIS_OPENAI_API_KEY` | Required by `evaluate_with_llm_judge` + `verify_citations` with `provider=openai` |
| `IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL` | Hard cost cap per LLM judge call (default `0.25`). `max_cost_usd` can lower it for a call, never raise it |
| `IRIS_RELEVANCE_JUDGE_MODEL` | A priced judge model id (e.g. `claude-haiku-4-5`). When set, with that provider's key, `answers_the_ask` asks this LLM judge on every evaluation that carries an input and gates on its relevance verdict — one judge call per evaluation, under the cost cap above and the two limits below. **Each such evaluation's input and output are sent to that model's provider (Anthropic or OpenAI) on your key**, with the personal data and credentials `no_pii` flags replaced first. Unset (the default), `answers_the_ask` reads the ask lexically and advises, and nothing is sent ([docs/llm-as-judge.md](https://github.com/iris-eval/mcp-server/blob/main/docs/llm-as-judge.md#the-relevance-judge-behind-answers_the_ask)) |
| `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD` | What the relevance judge may spend per UTC day, per tenant (default `1`). Kept in the database, so a restart does not reset it. A call is made only if its worst case fits in what is left; past that, `answers_the_ask` reads the ask lexically and `judge.withheld` is `daily_budget`. `0` stops every call |
| `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` | What every judge call together may spend per UTC day, per tenant (default `1`): the relevance judge, `evaluate_with_llm_judge` and `verify_citations` share it. Kept in the database, so a restart does not reset it. A call is made only if its worst case fits in what is left; past that, `answers_the_ask` reads the ask lexically (`judge.withheld: "daily_budget"`) and the judge tools answer `IRIS_BUDGET_EXCEEDED` (a citation, `daily_budget_reached`). `0` stops every call |
| `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD` | The name of `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` before 0.21.0, when it limited the relevance judge only. Read when the new name is unset, and it then limits every judge call; the server says so at startup |
| `IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST` | Relevance judge calls one request may make (default `20`): an OTLP batch or an `evaluate_runs` re-score judges its first 20 traces and reads the rest lexically, with `judge.withheld: "request_cap"` |
| `IRIS_RELEVANCE_JUDGE_REDACT` | `on` (default): every span `no_pii` flags (personal data and credentials) in the input and output is replaced by a `[REDACTED:<kind>#<n>]` marker before they are sent to the relevance judge. `off` sends them as they are |
| `IRIS_CITATION_ALLOW_FETCH` | Set to `1` to permit outbound HTTP in `verify_citations` (off by default). A tool argument can turn it off for a call, never on |
Expand Down
4 changes: 2 additions & 2 deletions claude-plugin/skills/iris-eval/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -237,8 +237,8 @@ result, including a passing one.
2. Put it in the environment of the process that runs Iris, not only your shell. Claude Code, Claude Desktop, Cursor and most MCP clients: the "env" block of the iris-eval entry in your MCP config — "iris-eval": { "command": "npx", "args": ["-y", "@iris-eval/mcp-server"], "env": { "IRIS_ANTHROPIC_API_KEY": "sk-ant-..." } } (IRIS_OPENAI_API_KEY for an OpenAI key). Docker: -e IRIS_ANTHROPIC_API_KEY=... on the run command. HTTP or CI: export it before starting iris-eval.
3. Restart the MCP session. A running process never sees a variable set after it started.
4. Confirm from inside your client: read iris://capabilities — judge.enabled must be true there. A key exported in your shell is not passed to the process your client spawns unless its config lists it. On a machine, `npx @iris-eval/mcp-server --self-test` prints the judge line for that shell, and GET /api/v1/health reports judge.enabled on a running dashboard.
5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Iris calls the provider directly with your key and never proxies it.
6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It spends at most IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD per UTC day (default 1 USD) and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why.
5. Spend guard: each call is capped by IRIS_LLM_JUDGE_MAX_COST_USD_PER_EVAL (default 0.25 USD) and refused before any spend if the worst case would exceed it. Every judge call Iris makes on your key draws on one daily budget, IRIS_LLM_JUDGE_DAILY_BUDGET_USD (default 1 USD per UTC day); past it, no call is made until midnight UTC. Iris calls the provider directly with your key and never proxies it.
6. Optional: set IRIS_RELEVANCE_JUDGE_MODEL to a priced model id (claude-haiku-4-5, for example) to have answers_the_ask ask the judge whether each answer addresses its ask, and fail an off-topic one. That is one judge call per evaluation that carries an input, on your key and under the cap above; the key alone never turns it on. Each call sends that input and output to the model's provider, with the personal data and credentials no_pii flags replaced first (IRIS_RELEVANCE_JUDGE_REDACT=off sends them as they are). It draws on that daily budget and makes at most IRIS_RELEVANCE_JUDGE_MAX_CALLS_PER_REQUEST calls per request (default 20); past either, answers_the_ask reads the ask lexically and says why.

## Example Workflows

Expand Down
Loading
Loading