Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions .claims.json
Original file line number Diff line number Diff line change
Expand Up @@ -5049,8 +5049,8 @@
],
"version": 1
},
"generatedAt": "2026-09-28T21:23:33.437Z",
"generatedFromCommit": "1b8c7fea",
"generatedAt": "2026-09-28T22:03:05.332Z",
"generatedFromCommit": "9852c053",
"generatorVersion": "1.0.0",
"llmJudgeTemplates": {
"count": 7,
Expand Down Expand Up @@ -9266,16 +9266,16 @@
"passed": null,
"total": 40
},
"totalCombined": 4174,
"totalCombined": 4182,
"vitestDashboard": {
"failed": 0,
"passed": 400,
"total": 400
},
"vitestRoot": {
"failed": 0,
"passed": 3734,
"total": 3734
"passed": 3742,
"total": 3742
}
},
"version": {
Expand Down
66 changes: 66 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -167,6 +167,72 @@ jobs:
- run: npm ci
- run: npx vitest run tests/unit/storage/migration-015.test.ts tests/unit/storage/trace-search.test.ts tests/unit/storage/search-query.test.ts tests/unit/tools/get-traces-search.test.ts

# The test jobs install better-sqlite3's prebuilt binary. When that download
# fails, `prebuild-install || node-gyp rebuild` compiles it here instead,
# against this Node's headers, and a binary compiled against Node 24.19+
# headers aborts the process the first time V8 frees one of its statements
# ("Assertion failed: (env) != nullptr", nodejs/node#65446; every 24.x
# release so far). A user whose download failed gets that binary too.
# This job compiles it on purpose on Linux, macOS and Windows and requires: the binary
# aborts under collection exactly when Iris predicts it; Iris never loads
# it there (it holds the store with node:sqlite, and says why); and the
# real server, started and stopped over stdio again and again, ends every
# session cleanly. Node 22 is the control: its headers never changed, so
# its binary is safe and Iris keeps the native driver. When a Node 24
# release carries the fix, the collect test here fails with the version
# to add to runtimeKeepsAddonHooks.
native-from-source:
name: native addon built from source (${{ matrix.os }}, Node ${{ matrix.node-version }})
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
node-version: [22, 24]
exclude:
# The node-gyp that ships with Node 22's npm 10 cannot identify the
# runner's Visual Studio 18 ("unknown version"), so nothing compiles
# there and npm drops the optional module (the install-without-
# better-sqlite3 job covers that path). The Node 22 control runs on
# Linux and macOS.
- os: windows-latest
node-version: 22
defaults:
run:
shell: bash
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: ${{ matrix.node-version }}
cache: npm
cache-dependency-path: package-lock.json
- name: Install, compiling better-sqlite3 here instead of downloading it
env:
npm_config_build_from_source: 'true'
# --foreground-scripts: the compiler's output is in the log, including when the optional build fails and npm drops the module.
run: npm ci --foreground-scripts
- name: The binary was compiled on this runner, against this Node's headers
run: |
# A prebuilt download holds only build/Release; node-gyp leaves its config beside it.
test -d node_modules/better-sqlite3 || { echo "::error::better-sqlite3 did not compile here and npm dropped it (it is optional); the install log above says why"; exit 1; }
test -f node_modules/better-sqlite3/build/config.gypi || { echo "::error::better-sqlite3 was not compiled here, so this job proves nothing"; exit 1; }
node -e "
const fs = require('fs');
const marked = fs.readFileSync('node_modules/better-sqlite3/build/Release/better_sqlite3.node').includes('RemoveEnvironmentCleanupHook');
console.log('Node ' + process.versions.node + ': binary carries the ObjectWrap cleanup-hook call: ' + marked);
if (process.versions.node.startsWith('24.') && !marked) { console.error('::error::Node 24 headers should have compiled the call in; this job proves nothing'); process.exit(1); }
"
- run: npx vitest run tests/integration/native-addon-collect.test.ts tests/integration/native-teardown-stdio.test.ts tests/unit/storage/driver.test.ts
- name: The self-test names the driver and why
run: |
IRIS_HOME="$RUNNER_TEMP/iris-home" IRIS_DB_PATH="$RUNNER_TEMP/iris-home/iris.db" node --import tsx src/index.ts --self-test | tee self-test.txt
if [ "${{ matrix.node-version }}" = 24 ]; then
grep -q "driver node: better-sqlite3 here was compiled against Node headers that abort" self-test.txt
else
grep -q "driver better-sqlite3" self-test.txt
fi

integration:
runs-on: ubuntu-latest
needs: test
Expand Down
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -112,6 +112,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **A failing test in the truthbase capture is reported as a failing test, by name, not as drift.** On `main` at 969719e the CI truthbase job said only "claims.json drifted from generator output". The tree was byte-identical to a pull-request head whose same job had passed; one test had failed during the capture (3,729 of 3,730), was recorded as `failed: 1`, and no line of the log named it. The capture now refuses a run with a failing test and lists each one with its file and the first line of its failure, and `claims:check` names every field that differs from the generator's output (for that run it would have printed `tests.vitestRoot.passed: committed 3730, generated 3729`).
- **The Python client's publish workflow parses again, and every workflow is linted on every pull request.** Since the hash-pinned build landed, `publish-python.yml` held a plain YAML value with `:all: -r` in it, which YAML reads as a mapping, so GitHub refused the file on every push. Nothing noticed, because the workflow ran only on a `py-v*` tag. The step is now a block scalar. A new CI job runs actionlint 1.7.12, from an image pinned by digest, over every workflow in `.github/workflows`, including each `run:` script through shellcheck. It reported the parse error on the old file, and 6 shellcheck findings elsewhere, now fixed (four unguarded `ls *.tgz` globs in `ci.yml`, an unused loop variable and an unquoted `kill $(cat …)` in `lighthouse.yml`). `publish-python.yml` now also runs its build job on every pull request (the hashed install, the build and the fresh-environment import), and publishes only for a tag.
- **`o1-mini` is priced at $1.10 in and $4.40 out per million tokens, not $3 and $12.** The judge's table carried the model's launch price; OpenAI's model page (developers.openai.com/api/docs/models/o1-mini) lists $1.10 / $4.40, and its pricing page no longer lists the model. The judge's cost cap and cost report for `o1-mini` were 2.7 times too high, so the cap refused calls that fit. Every other row was checked against claude.com/pricing and developers.openai.com/api/docs/pricing on 2026-09-28 and matches, and the table's read date is now 2026-09-28.
- **On Node 24, Iris no longer loads a better-sqlite3 that would abort the server when it frees a statement (#719).** Node 24.19.0 changed the `node::ObjectWrap` header without the runtime change it needs ([nodejs/node#65446](https://github.com/nodejs/node/issues/65446)). A better-sqlite3 compiled against that header, as npm does when the prebuilt download fails, aborts the process with `Assertion failed: (env) != nullptr` the first time V8 collects one of its statements, whether the database is open or closed. On Node 24.21.0 with 12.11.1, measured: a binary built from source aborted in 10 runs of 10, the prebuilt binary in 0 of 10. It aborted CI once, in 1 of the 62 failed jobs of the last 400 runs, the one whose install had compiled the addon. Iris now reads the binary before loading it. When the binary carries the new call and the runtime lacks the fix (every 24.x release so far, and 26.x before 26.4.0), Iris holds the store with Node's built-in SQLite and says why on stderr and in `--self-test`; `GET /health` reports `driver: "node"`. `IRIS_SQLITE_DRIVER=native` refuses instead. `npm rebuild better-sqlite3` restores the prebuilt binary, which is safe. A new CI job compiles the addon from source on Node 24 on Linux, macOS and Windows, and on Node 22 on Linux and macOS. It requires the binary to abort under collection exactly when Iris predicts it, and a real server started and stopped over stdio six times to end every session cleanly. With the check turned off, that stdio test failed 5 runs of 5 on the compiled binary.
- **A stdio server shuts down in order when the client closes its stdin.** MCP clients end a stdio session that way. The server used to drain and exit without closing the store, and a search-index build kept running for the client that had gone: 1.5 s on 80,000 traces, against 32 ms now. With the dashboard running, the end of stdin still does not stop the process.
- **The published test counts are no longer recorded from a run in which a test file failed to load.** `npm run claims:capture-tests` read only vitest's totals, and a file that does not parse or throws at import counts none of its tests and no failure, so a run eleven tests short once recorded every test passing. The capture now refuses such a run, naming each file and the first line of its error, and writes nothing; it also refuses a run vitest marks failed with no failing test, and a run with a skipped test (which it used to fall back from to the committed counts). `--report root=<file>` and `--report dashboard=<file>` read an existing vitest JSON report instead of running the suite, which is how `tests/unit/scripts/capture-report.test.ts` proves the refusal on a real report of a file that fails to parse and one that throws at import.
- **The Python recorder batches as it says it does.** `IrisRecorder` waits `flush_interval` (0.25 s by default) to gather traces into one request, but every new trace woke that wait early, so an application recording steadily sent one request per trace and a full queue was never reached. The interval is now a deadline that only `flush()` or `close()` ends early. `test_the_queue_is_bounded_and_drops_the_oldest` yields between records, and fails on the old recorder every run.
- **The Failures and Moments pages no longer re-run the regression-alarm watcher once per trace (#680).** To find the alarms each trace in its window carries, the route re-ran every rule's CUSUM watcher over the agent's log up to that trace, so the cost grew with the square of the window and with the number of rules. The watchers now run once per request, and once more for each size of the rule family the log grows through; a log whose rules and runs are there from the start takes one pass. Every trace gets exactly the alarms it got before: a test checks the one pass against the per-trace answer for every trace of seeded logs where rules and runs start part-way through, where traces share a timestamp, and where alarms fire. On one agent's 500 traces the Failures and ranked Moments responses are byte-identical before and after. Measured on that window with no other load: with 25 rules, a warm request takes 286–344 ms, down from 2.5–2.6 s, and the first request after a start 4.5 s, down from 6.8 s; with 3 rules, 33–45 ms, down from 94–101 ms. The first-request cost that remains is the simulation that draws each stream's alarm line, which is memoised once drawn.
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -467,7 +467,7 @@ Every variable `--help` documents. CLI flags take precedence over environment va
| `IRIS_PORT` | HTTP transport port (1-65535, default `3000`) |
| `IRIS_HOME` | Directory for all per-user files: `config.json`, `iris.db`, `custom-rules.json`, `audit.log`, `preferences.json` (default `~/.iris`) |
| `IRIS_DB_PATH` | SQLite database path (overrides `IRIS_HOME` for the DB only) |
| `IRIS_SQLITE_DRIVER` | Which SQLite driver holds the database: `native` (better-sqlite3, the default) or `node` (Node's built-in `node:sqlite`, Node 22.13+). Unset: native, and when the native module cannot load Iris warns once and falls back to the built-in |
| `IRIS_SQLITE_DRIVER` | Which SQLite driver holds the database: `native` (better-sqlite3, the default) or `node` (Node's built-in `node:sqlite`, Node 22.13+). Unset: native, and when the native module cannot load (or is a build that would abort on this Node) Iris warns once and falls back to the built-in |
| `IRIS_SEARCH_BUDGET_MS` | How long one trace search (`q`) may read before it answers with the matches it found so far and `search.complete: false`, in milliseconds (50 to 60000, default `1000`). A search holds other requests while it reads, so this is also the longest it can make them wait. Also `storage.searchBudgetMs` in `config.json` |
| `IRIS_LOG_LEVEL` | Log level: `debug`, `info`, `warn`, `error` |
| `IRIS_DASHBOARD` | `true`/`1`/`yes`/`on` enables the web dashboard; `false`/`0`/`no`/`off` disables it (also overrides `dashboard.enabled` in `config.json`) |
Expand Down Expand Up @@ -602,7 +602,7 @@ npm update -g @iris-eval/mcp-server

**On a platform with no prebuilt `better-sqlite3`, the install still succeeds.** `better-sqlite3` is an optional dependency: when npm can neither download a prebuilt binary for your Node and platform nor compile one (compiling needs Python and a C++ toolchain — Visual Studio's C++ build tools on Windows), npm prints the build error, skips the module, and finishes the install. Iris then runs on Node's built-in SQLite, and says so: startup prints one line on stderr naming why, and `--self-test` shows `driver node: better-sqlite3 is not installed …`. To get the native driver back, install it where a prebuild or a toolchain exists (`npm install better-sqlite3` in the project; for a global install, install Iris again with `npm install -g @iris-eval/mcp-server` once a toolchain is available). CI installs the packed server with the native build forced to fail on every change, and requires the install to finish and the self-test to store and read a trace on the built-in.

Iris keeps everything in one SQLite file, opened by `better-sqlite3` — a native addon that is downloaded or compiled for your Node and platform. **When that module cannot load, Iris falls back to Node's built-in SQLite** (`node:sqlite`, Node 22.13 or later) with one warning on stderr, so a missing prebuild is a slower start rather than a dead one; `IRIS_SQLITE_DRIVER=node` chooses the built-in on purpose, `native` forbids the fallback. The built-in is opened with extension loading off and `trusted_schema` off; Node prints its own `ExperimentalWarning: SQLite is an experimental feature` line on stderr when it loads, and Iris does not silence it. `--self-test` and `GET /health` name the driver in use; every number on the proof page was measured on the native driver, and the test suite runs on both in CI.
Iris keeps everything in one SQLite file, opened by `better-sqlite3` — a native addon that is downloaded or compiled for your Node and platform. **When that module cannot load, Iris falls back to Node's built-in SQLite** (`node:sqlite`, Node 22.13 or later) with one warning on stderr, so a missing prebuild is a slower start rather than a dead one. It does the same, before loading it, for a `better-sqlite3` compiled on your machine against Node 24.19 or later headers: on every 24.x release so far such a binary aborts the whole process the first time it frees a statement (`Assertion failed: (env) != nullptr`, [nodejs/node#65446](https://github.com/nodejs/node/issues/65446)), and `npm rebuild better-sqlite3` replaces it with the prebuilt binary, which is safe. `IRIS_SQLITE_DRIVER=node` chooses the built-in on purpose, `native` forbids the fallback. The built-in is opened with extension loading off and `trusted_schema` off; Node prints its own `ExperimentalWarning: SQLite is an experimental feature` line on stderr when it loads, and Iris does not silence it. `--self-test` and `GET /health` name the driver in use; every number on the proof page was measured on the native driver, and the test suite runs on both in CI.

### Node.js version

Expand Down
2 changes: 1 addition & 1 deletion docs/api-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -1444,7 +1444,7 @@ The one health contract (0.15.0). Unauthenticated by design — no key, no sessi
}
```

- `driver` — the SQLite driver behind the store: `better-sqlite3` (the native addon, the default) or `node` (Node's built-in `node:sqlite`, chosen with `IRIS_SQLITE_DRIVER=node` or fallen back to when the native module cannot load); `null` on a transport started without storage.
- `driver` — the SQLite driver behind the store: `better-sqlite3` (the native addon, the default) or `node` (Node's built-in `node:sqlite`, chosen with `IRIS_SQLITE_DRIVER=node` or fallen back to when the native module cannot load or is a build that would abort on this Node); `null` on a transport started without storage.
- `search_worker` — where searches run: `ready` (on their own thread, so a slow search never holds other requests), `not_started` (the thread starts with the first search), `unavailable` (it could not start on this machine, so searches run on the main thread; `detail` gives the reason, and the server also logs it once), or `not_used` (a store in memory, which searches on the main thread by design); `null` without storage. It is informational: search works in every case, so it never makes `status` `degraded`. `--self-test` prints the same line.
- `checks.storage` — the database answered a count. The count itself is not reported: this endpoint answers without a key, so it says whether the store works, not how much it holds (the number is `total` on the authenticated `GET /api/v1/traces`); `checks.rules_store` — the deployed custom-rules file reads and parses; `checks.migrations` — every migration this build knows is applied, with the numbers so a schema that is behind is visible before a query fails. Each is `ok`, `fail`, or `absent` when there was nothing to check.
- `status` is `ok` only when no check failed; otherwise `degraded`, with HTTP **503**, so a probe that reads only the status code is right.
Expand Down
2 changes: 1 addition & 1 deletion docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -216,7 +216,7 @@ dashboard/ React SPA (separate Vite build)

1. `index.ts` parses CLI args with `node:util.parseArgs`.
2. `loadConfig()` reads `~/.iris/config.json` and validates it against a strict schema (`src/config/schema.ts`) — a key Iris does not read, or a value of the wrong type, refuses startup naming it — then layers the file, the `IRIS_*` env vars and the CLI args over the defaults, in that order.
3. `createStorage()` instantiates `SqliteAdapter`, which opens the file through the driver seam (`src/storage/driver.ts`: `better-sqlite3` by default, Node's built-in `node:sqlite` when chosen with `IRIS_SQLITE_DRIVER=node` or when the native module cannot load) and calls `initialize()` to set the busy timeout, enable WAL mode, turn on foreign keys and run pending migrations — every migration is typed on the seam, not on a driver.
3. `createStorage()` instantiates `SqliteAdapter`, which opens the file through the driver seam (`src/storage/driver.ts`: `better-sqlite3` by default, Node's built-in `node:sqlite` when chosen with `IRIS_SQLITE_DRIVER=node` or when the native module cannot load or is a build that would abort on this Node) and calls `initialize()` to set the busy timeout, enable WAL mode, turn on foreign keys and run pending migrations — every migration is typed on the seam, not on a driver.
4. `createIrisServer()` creates the MCP `McpServer` instance, instantiates `EvalEngine` with the configured threshold, and registers all tools and resources.
5. Based on `config.transport.type`:
- **stdio**: Creates `StdioServerTransport` and connects.
Expand Down
Loading
Loading