diff --git a/CHANGELOG.md b/CHANGELOG.md index ce802a39..f9914e45 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -28,7 +28,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **BREAKING — A tool argument can no longer widen what the operator allows for fetching or spend.** Three arguments could. `verify_citations` with `allow_fetch: true` fetched the cited URLs on a server whose operator had not set `IRIS_CITATION_ALLOW_FETCH=1`; its `domain_allowlist` was added to the operator's `IRIS_CITATION_DOMAINS` instead of narrowing it; and `max_cost_usd` on `evaluate_with_llm_judge` and `max_cost_usd_total` on `verify_citations` replaced the operator's cap with any larger number. An agent's arguments can be steered by text it read, including text stored in Iris, so each is now a ceiling: an argument can narrow the operator's setting for one call and never widen it. When one asks for more, the operator's setting applies and the call goes on, the response carries a warning with code `IRIS_ARGUMENT_NARROWED` naming the argument and the setting that applied, and the stored evaluation keeps the same fact (`provenance.narrowed`) with a sentence to the operator in its `interpretations`, on every later read. Fetching now needs `IRIS_CITATION_ALLOW_FETCH=1` on the server; a caller that relied on `allow_fetch: true` gets no fetch and the warning. The per-call cap on `verify_citations` is a new setting, `IRIS_CITATION_MAX_COST_USD_TOTAL`, default 1 USD as before. - **BREAKING — Every judge call on your key draws on one daily budget.** `evaluate_with_llm_judge` and `verify_citations` had a per-call cap and no daily limit, so an agent calling either in a loop, or steered into doing so, could spend the key without end; only the relevance judge had a daily budget. All three now share one, `IRIS_LLM_JUDGE_DAILY_BUDGET_USD` (default 1 USD per UTC day, per tenant, kept in the database). A call is made only if its worst case fits in what is left. When it does not, nothing is spent: `evaluate_with_llm_judge` answers `IRIS_BUDGET_EXCEEDED` (retryable, with the time the budget resets), and `verify_citations` marks the citation `daily_budget_reached` and makes no further call. A deployment that spends more than 1 USD a day on the judge tools raises the budget. The variable was `IRIS_RELEVANCE_JUDGE_DAILY_BUDGET_USD`; that name is still read when the new one is unset, and the server says so at startup. In the Claude Desktop extension the setting is now "LLM judge daily budget", and a value set under the old one is not carried over. Today's spend is `judge.dailyBudget` on `iris://capabilities`. - **BREAKING — Stored text comes back fenced, and a page of traces carries snippets.** Iris hands back what agents, their users and their tools wrote, and the agent reading it sits beside tools that delete traces and rules; a sentence planted in a trace reached the model as plain JSON. Every read that returns stored text now puts each value someone else wrote inside a tag carrying an id made for that response, `…`, the fence Iris already used for its own judge, and the response starts with `untrusted: { id, notice }`. The fence is on the values, so it reaches the model from the text block, from `structuredContent` and from a resource alike. A short value with no whitespace (an agent name, a session id, a tool name) is left as it is, so it still works as a filter. Covered: `get_traces`, `list_rules`, `iris://traces/{trace_id}`, `iris://evaluations/{id}` (a custom rule's name and a judge's rationale; what Iris wrote stays outside), `iris://audit` and `iris://dashboard/summary`. `get_traces` cuts each fenced value to 500 characters unless `include_text: true`, and says what it cut in each trace's `cut`. `deploy_rule` and `evaluate_output`'s `custom_rules` refuse a value carrying a fence tag, and `deploy_rule` takes the rule names the dashboard always required (letters, digits, dot, dash, underscore). `get_traces` and `list_rules` advertise `openWorldHint: true`. A script that read a trace's text from `get_traces` reads it inside the tags, and passes `include_text: true` for more than 500 characters. -- **Two dependency advisories patched.** `proxy-addr` 2.0.8 ([GHSA-jqcg-44mw-7w3h](https://github.com/advisories/GHSA-jqcg-44mw-7w3h), critical): a client could pass as a trusted proxy through an IPv4-mapped IPv6 address. The check it fixes runs only when Express's `trust proxy` is set to an address range, and Iris leaves `trust proxy` off. `source-map-js` 1.2.2 ([GHSA-68fv-2mgg-jv7q](https://github.com/advisories/GHSA-68fv-2mgg-jv7q), high) is in the build and test tools only and does not ship in the package. +- **Three dependency advisories, each patched within a day of GitHub publishing it; none affects Iris as shipped.** `proxy-addr` 2.0.8 ([GHSA-jqcg-44mw-7w3h](https://github.com/advisories/GHSA-jqcg-44mw-7w3h), critical) fixes how a trusted proxy address range is matched; that matching runs only when Express's `trust proxy` is set to a range, and Iris leaves `trust proxy` off. `source-map-js` 1.2.2 ([GHSA-68fv-2mgg-jv7q](https://github.com/advisories/GHSA-68fv-2mgg-jv7q), high) is in the build and test tools only. `@modelcontextprotocol/client` 2.3.1 ([GHSA-6qxp-vccf-f47h](https://github.com/advisories/GHSA-6qxp-vccf-f47h), high) is in the OpenTelemetry example only. ## [0.20.0] - 2026-10-03