chore(eval): integrate Retrieval Eval v2 and close Phase 1 - #215
Merged
Merged
Conversation
Two gates the v1.5 eval path needs to be trustworthy on every platform. **`npm run eval` could not start on Windows.** The three scripts resolve `node_modules/.bin/electron.cmd` and spawn it. Node 24 refuses to spawn a `.cmd`/`.bat` without `shell: true` and fails with `EINVAL`, so the harness was unrunnable on the platform this project is developed on. The `electron` package exports the path to the real executable the wrapper runs, which spawns directly on every platform and needs no shell. **A Windows run rewrote the committed baseline.** `path.relative` returns backslashes on Windows, so `config.corpus` was written as `eval\corpus` instead of `eval/corpus`. The file's whole contract is that it is identical on every machine — the CI determinism check diffs it — and a Windows run silently broke that. The label is now normalised to POSIX separators. Verified on Windows with the pinned model: `npm run eval` runs and leaves `docs/eval/baseline-v1.5.json` byte-identical to the committed file.
… config The harness measured a retriever nobody runs. Production retrieved at `topK: 3` with `threshold: 0.5`; the harness ran `topK: 10` with `threshold: 0` and reported `Recall@5 = 1.0000` for a pipeline that silently drops rank 4 at cosine 0.47. The first child of #192 asks for the two to be the same configuration. **One K was doing two jobs.** `RetrievalRequest.topK` was both the first-stage width (KNN neighbours, BM25 limit) and the number of passages delivered. In hybrid that made fusion nearly a no-op: dense contributed `topK`, BM25 contributed `topK`, RRF fused at most `2 * topK`, and the result was sliced straight back to `topK`. The two stages are now named: ``` RetrievalRequest: candidateK first-stage width per channel (default 20) topK passages delivered (default 5) ``` `effectiveCandidateK` enforces `candidateK >= topK`, so a caller that asks for more results than the default pool (the search palette, MCP) is never silently capped. `HybridRetriever` fuses the wide pool and truncates once, after fusion. `DenseRetriever` queries `candidateK` and slices to `topK` — equivalent to before for a single strategy, since a prefix of a ranking is the same ranking. The trace and the #157 snapshot gained `candidateK`. A snapshot written before this change backfills it from `topK`, which is what that retrieval actually did. **The harness now defaults to production** (`candidateK: 20`, `contextK: 3`, `threshold: 0.5`), with `--eval-candidate-k=` / `--eval-context-k=` / `--eval-threshold=` to move them deliberately. Ranking metrics are computed at `candidateK` depth, not `contextK`: `Recall@10` needs ten results, and truncation only takes a prefix, so it cannot change the ranking being measured. `contextK` is recorded so the report describes the whole online path. `docs/eval/baseline-v1.6.{json,md}` is the new frozen baseline; v1.5 is kept as history for the #77/#78 deltas, and the CI determinism check moves to v1.6. ## What this did and did not change On the current 19-chunk corpus the production configuration produces the **same** metrics as v1.5 — `threshold: 0.5` is non-binding and `candidateK: 20` exceeds the corpus, so nothing is filtered and no ranking changes. That is the corpus limitation #192 describes, not a result. What changed is that the harness now *states* the production parameters instead of assuming different ones. Hybrid still leads dense on Recall@1 (0.8667 vs 0.8333), MRR (0.9444 vs 0.9278) and nDCG@10 (0.9561 vs 0.9437) with a wide pool; dense stays the default because the adoption rule still cannot move on a saturated Recall@5. ## Testing - `npm run typecheck` — clean - `npm test` — 481 pass (3 new: trace keeps the two Ks apart, the width invariant, the pre-#77 snapshot backfill) - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` - `npm run eval:retrieval` — dense/sparse/hybrid unchanged in ordering Part of #192 (child 1).
Recall@K, MRR and nDCG@10 could not express three things the epic needs, and a single aggregate average was hiding a gap the corpus already contained. Child 3 of #192. **Two metrics added.** `hitRate@5` is whether *any* of the first five passages covers ground truth; `MAP@10` combines ranking position with coverage, so pulling a second relevant passage from rank 9 to rank 2 moves it while Recall@5 sits still. Hit rate is deliberately blunt next to Recall@5: a two-passage question that finds one scores 1.0 and 0.5 respectively, and both facts matter — "the model had a chance" is not "the material was complete". **Every metric is now reported per query type.** `questions.jsonl` gained an optional `type`, the harness groups by it, and the report renders a table. The aggregate was already concealing something: | Type | n | Recall@5 | nDCG@10 | Hit rate@5 | MAP@10 | | --- | --- | --- | --- | --- | --- | | cross-lingual | 1 | 1.0000 | **0.6309** | 1.0000 | **0.5000** | | exact | 6 | 1.0000 | 1.0000 | 1.0000 | 1.0000 | | multi-hop | 2 | 1.0000 | 0.9599 | 1.0000 | 0.9167 | | semantic | 18 | 1.0000 | 0.9312 | 1.0000 | 0.9074 | | zh | 3 | 1.0000 | 1.0000 | 1.0000 | 1.0000 | | **all** | 30 | 1.0000 | 0.9437 | 1.0000 | 0.9222 | The one cross-lingual question (`q030`, Chinese over an English source) ranks far worse than everything else. Recall@5 = 1.0000 reported that as a success; nDCG@10 and MAP@10 are what make the multilingual gap visible. That is the metric doing its job on the existing 30 questions, before the corpus grows. Untagged questions group under `untagged` rather than being dropped, and the types in use are documented in `eval/README.md`. ## Testing - `npm run typecheck` — clean - `npm test` — 485 pass, 4 new: hit rate vs recall, hit rate@k bounds, AP position sensitivity, AP's repeat-counts-once rule - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` Part of #192 (child 3).
…test `threshold: 0.5` was hand-picked, and a cosine score has no universal meaning — its distribution depends on the embedding model, the language, the query type and the chunk length. Child 4 of #192. **A deterministic split, owned by the harness.** `--eval-split=validation|test` partitions the questions by a hash of the id, so the same `questions.jsonl` cuts the same way on every machine and both arms go through the same code path that produces the frozen baseline. A parameter chosen on the questions it is scored on is fitted, not measured. **`npm run eval:threshold`** sweeps `0 / 0.3 / 0.4 / 0.5 / 0.6`, selects on `validation`, and reports the winner on `test`. It also counts a metric the baseline does not carry: the **no-result rate**. A higher threshold can look better on a ranking metric while quietly making the product answer "not in your sources" more often, and that trade is invisible unless it is counted. ## The result here is a non-result, and it is reported as one | Threshold | Recall@5 (val) | nDCG@10 (val) | No-result (val) | nDCG@10 (test) | No-result (test) | | --- | --- | --- | --- | --- | --- | | 0 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.3 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.4 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.5 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.6 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | The sweep is **flat**: every threshold produces identical metrics and never filters a passage. E5 does not score these query/chunk pairs below 0.6, so on this corpus the threshold is non-binding. The report says **"no evidence to change `threshold = 0.5`"** rather than nominating the tie-break winner, because moving a product parameter on a flat sweep would be noise dressed as a result. The tie-break rule (widest threshold among equals) is stated so it can be argued with. That is the honest outcome, and it is another instance of the corpus limitation #192 describes. The machinery is what ships here; the number is what the corpus cannot yet support. ## Testing - `npm run typecheck` — clean - `npm test` — 489 pass, 4 new: split bounds, split determinism, validation/test partition without overlap, `all` preserves order - `npm run eval` — `docs/eval/baseline-v1.6.json` regenerated with `config.split` - `npm run eval:threshold` — reruns cleanly and rewrites the same tables Part of #192 (child 4).
The v1.5 rule was "Recall@5 must improve and nDCG@10 must not regress". On a corpus where dense already scores Recall@5 = 1.0000, no strategy can improve Recall@5, so the rule was not strict — it was **unsatisfiable**. Every comparison came back "inconclusive", including a hybrid that was better on Recall@1, MRR and nDCG@10. The default never moved, not because hybrid lost but because the rule could not return a verdict. Child 10 of #192. **The amendment.** A metric at its maximum has no headroom and is not allowed to decide. The deciding metric is the first one with headroom, in the order `recallAt5`, `ndcgAt10`, `mrr`, `mapAt10`; a strategy clears the rule when it improves that metric and regresses none of the others. Saturation is detected and reported rather than silently blocking every change. **The rule is code, not prose.** It lives in `src/main/eval/adoption.ts` with 8 unit tests, because it decides whether a shipped default moves and testing it by reading the sentence the script prints would only test the sentence. The experiment script imports it, so there is one definition. ## The v1.5 stalemate resolves ``` | Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | MAP@10 | Query p95 | | dense (vector) | 0.8333 | 1.0000 | 0.9278 | 0.9437 | 0.9222 | 14.06 ms | | sparse (BM25) | 0.7333 | 0.8667 | 0.8056 | 0.8184 | 0.8000 | 2.63 ms | | hybrid (RRF dense + BM25) | 0.8667 | 1.0000 | 0.9444 | 0.9561 | 0.9389 | 22.32 ms | ``` `recallAt5` is reported as saturated; the deciding metric is `ndcgAt10`; **hybrid clears the rule** and regresses none of the four metrics, at a p95 cost of ~8 ms. The script reports the measurement and does not flip the default — changing the shipped strategy is a product decision, and it is stated that way in the output rather than implied by a green checkmark. `docs/eval/retrieval-v1.6.{json,md}` is the regenerated experiment against the v1.6 baseline; `retrieval-v1.5.*` is kept as the record of the superseded rule. ## Testing - `npm run typecheck` — clean - `npm test` — 497 pass, 8 new: saturation detection, deciding-metric priority, the float-slack regression check, the resolved stalemate, a trade that regresses another metric being refused, the baseline not clearing against itself, all-saturated, and best-candidate selection - `npm run eval` + `npm run eval:retrieval` — regenerated and re-run Part of #192 (child 10).
…odel The harness measured the retriever but never the window the prompt actually gets. `evidenceK: 5` was a separate constant from the `contextK: 3` production uses, so the one metric that looked at a window looked at a different one than the product does. Child 5 of #192. **`contextK` is now the window for both context metrics.** - `contextPrecision@contextK` — of the first `contextK` passages, the share covering ground truth. This is what `evidencePrecisionAt5` was, at the production width. - `contextRecall@contextK` — the share of needed ground-truth blocks that made it into that window. Distinct from `Recall@10`: a block found at rank 4 is invisible when `contextK = 3`, and that is a product fact, not a ranking fact. Both are **deterministic**: the dataset says which blocks answer the question, so no model is needed to score a window. Together they are the trade-off a `contextK` decision makes — a wider window finds more and carries more noise — which is what the sweep in the next child needs. Baseline: `contextPrecision@3 = 0.3556`, `contextRecall@3 = 1.0000` — the needed evidence is always inside the top 3 on this corpus, but only about a third of what is inside the window is relevant. That second number is the one that says the window is paying for passages that do not answer the question. **What is deliberately not here.** Faithfulness, completeness, answer correctness and noise sensitivity need a generative model. The harness runs offline with only the pinned embedding model — the same constraint that keeps the reranker unmeasured (#170) — so this PR does not add them and does not fake them: the report's Definitions section now says so, and the "Not evaluated" line in `eval-retrieval.mjs` points at the same constraint. Adding an LLM judge is a separate change that has to solve model pinning first, not a line of code. `evidenceK` is removed from the harness config, so there is exactly one context width. ## Testing - `npm run typecheck` — clean - `npm test` — 497 pass - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` - `npm run eval:retrieval` — regenerated; hybrid still clears the amended rule Part of #192 (child 5, deterministic half).
The epic asks for a grid over `candidateK` and `contextK` with latency, index size and context size recorded next to quality, and explicitly not for a single aggregate "RAG score". Child 7 of #192. `npm run eval:sweep` runs the real harness over `strategy × candidateK {5,10,20,40} × contextK {3,5,8}` (24 runs, chunking fixed) and writes `docs/eval/sweep-v1.6.{json,md}`. The harness gained `contextChars` per question — the size of the context window, reported as **characters, not tokens**, because the harness pins an embedding model and no generation tokenizer. ## What the dashboard already shows Two of the three axes are decided by this corpus, and one of them decisively: - **`candidateK` changes nothing.** Every metric is identical from 5 to 40, because the corpus is 19 chunks and the relevant passages are already inside the top 5. This is the saturation problem #192 describes, now visible on the axis it affects. - **`contextK` has a clear optimum here: 3.** Context recall is 1.0000 at every width, while precision falls `0.3556 → 0.2133 → 0.1333` and the window grows `2079 → 3327 → 5312` characters as it widens. Wider adds prompt cost and noise for **no** recall. The production default is already 3, and this is the first evidence that it is the right 3 rather than a guess. - **hybrid beats dense** on nDCG@10 (0.9561 vs 0.9437) and MAP@10 (0.9389 vs 0.9222) at every setting, with no metric regressing — consistent with #197. The report labels the grid maximum as **not** a recommendation, because selecting on the same questions is how a benchmark becomes a lookup table; the adoption rule and the `validation`/`test` split are what keep the decision honest. Timing p95 is reported per row but is noisy at this sample size; it is in the table so a latency cost can be seen, not so it can be ranked. ## Testing - `npm run typecheck` — clean - `npm test` — 497 pass - `npm run eval` — regenerated (`contextChars` added to `perQuestion`) - `npm run eval:sweep` — 24 runs, dashboard written Part of #192 (child 7).
The sweep contained cells the harness cannot fill. The harness retrieves `candidateK` passages and the context metrics look at the first `contextK` of them, so `candidateK=5, contextK=8` reports on five passages while claiming eight. The dashboard showed `contextK=5` and `contextK=8` at `candidateK=5` as **identical** — and because `evidencePrecisionAtK` divides by the passages actually retrieved, nothing exposed it. Found in the #192 review. A wrong number that looks like a measurement is worse than a failure, so the invariant is enforced rather than documented: - **The harness refuses it.** `contextK > candidateK` throws before indexing, naming both numbers: `contextK (8) cannot exceed candidateK (5): the harness retrieves candidateK passages, so a wider context window can never be filled.` - **The sweep skips those cells** and says so. The report lists the skipped combinations in a "Skipped cells" section instead of quietly omitting rows, because "we did not measure this" and "this measured the same as its neighbour" are different statements and the old output showed the second while meaning the first. `docs/eval/sweep-v1.6.*` is regenerated: 22 rows instead of 24, and the misleading `candidateK=5 / contextK=8` row is gone rather than silently equal to `5`. The production cells (`candidateK=20, contextK ∈ {3,5,8}`) are unaffected. ## Testing - `npm run typecheck` — clean - `npm test` — 502 pass, including 3 new harness-invariant tests that need no vector store, because the check runs before any database work: the refusal, the message naming both numbers, and `contextK == candidateK` being accepted - `node_modules/electron/dist/electron.exe . --eval-harness --eval-candidate-k=5 --eval-context-k=8` — refused with the message above - `npm run eval:sweep` — 22 rows, 2 skipped and listed Part of #192 (review follow-up).
The corpus contract corpus v2 needs. Without it, two of the four things #192 asks the dataset to cover cannot be represented at all: a question whose right answer is "your sources do not say", and a validation/test assignment that is chosen rather than computed. **Unanswerable questions.** `answerable: false` with `relevant: []`. They are excluded from `metrics` and `byType` — `recallAtK` treats no ground truth as 0/0, and averaging "correctly refused" into "missed" would corrupt the table — and reported as their own group: `noResultCount`, `noResultRate` and `meanRetrieved`. On this row **higher no-result is better**, which is the opposite of how it reads everywhere else. `assertQuestionShape` refuses the reverse combination in either direction, because both are silent: an answerable question with no ground truth reads as a permanent miss, and an unanswerable one carrying ground truth reads as a normal hit. **An explicit split manifest.** `eval/splits.json` replaces the hash of the question id. A hash is reproducible but not *stable*: adding a question moved others between the sides, and a rare query type could end up entirely on one side without anyone choosing that — which is what had happened (`multi-hop` and `cross-lingual` were both entirely on `test`). A question with no manifest entry is now **refused** rather than defaulted, so a new question cannot leak into the reporting side. **Six unanswerable questions** are added (three topic-adjacent, three plainly out-of-scope). Their specifics were checked absent from the corpus before being written, so they are genuinely unanswerable rather than believed to be. ## What the first measurement shows ``` unanswerable: { questions: 6, noResultRate: 0, meanRetrieved: 19 } ``` Every unanswerable question returns **the entire corpus** — 19 of 19 chunks — at `threshold: 0.5`, and the threshold sweep is flat all the way to `0.6`: refusal stays `0/6` and `meanRetrieved` stays at the cap. So on this corpus the threshold cannot separate a relevant passage from an unrelated one, and "Who won the 2018 FIFA World Cup?" is answered with 19 passages of river monitoring, tidal energy and tea storage. That is the evidence the threshold decision did not have before, and it is also the corpus limitation #192 describes: it does not say `0.5` is wrong, it says this corpus cannot tell. The reports say so rather than recommending a change. Retrieval metrics are unchanged (`recallAt5` 1.0000), because the six new questions are correctly kept out of them. ## Testing - `npm run typecheck` — clean - `npm test` — 504 pass, including the manifest contract (unknown side rejected, missing entry refused, `all` needs no manifest), the answerable/unanswerable tie, and the harness's window invariant - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` - `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated; the threshold rule now gates on answerable quality and then maximises unanswerable refusal ## Scope This is the **contract**, plus the smallest honest corpus increment that exercises it. Hard negatives and near-duplicate documents — the thing that would bring `Recall@5` off 1.0000 and give `candidateK` something to do — are the next step, not this PR. Part of #192 (child 2, stage 1 of the corpus).
…saturated The corpus had exactly **one** cross-lingual question (`q030`, Chinese over an English source), which is not enough to call anything a weakness — correctly flagged in the #192 review as a signal rather than a finding. This adds eight more Chinese questions over **English** sources, reusing the already-verified ground truth of their English counterparts: the fact is the same, only the query language changes, which is precisely what the multilingual model has to bridge. Cross-lingual goes from n=1 to n=9. ## Recall@5 is no longer 1.0000 ``` before after Recall@5 1.0000 0.9211 nDCG@10 0.9437 0.8476 MAP@10 0.9222 0.7965 hitRate@5 1.0000 0.9211 ``` The benchmark was saturated on this corpus, which is why no retrieval change could ever clear the adoption rule and why `candidateK` appeared to do nothing. Eight questions were enough to bring Recall@5 off the ceiling. ## The finding, at a size that supports it | type | n | Recall@5 | nDCG@10 | MAP@10 | | --- | --- | --- | --- | --- | | **cross-lingual** | **9** | **0.6667** | **0.5032** | **0.3446** | | exact | 6 | 1.0000 | 1.0000 | 1.0000 | | multi-hop | 2 | 1.0000 | 0.9599 | 0.9167 | | semantic | 18 | 1.0000 | 0.9312 | 0.9074 | | zh | 3 | 1.0000 | 1.0000 | 1.0000 | Chinese questions over a **Chinese** source (`zh`) still score 1.0000. Chinese questions over an **English** source score 0.6667. So the gap is specifically **cross-lingual retrieval**, not Chinese, and it is now measured on nine questions rather than asserted from one. This is the kind of finding that was not defensible before: `q030` alone gave nDCG 0.63 and the honest conclusion was "a signal worth testing". With n=9 the same direction holds at the same magnitude, so it is a real weakness of the current pipeline on this corpus. ## What it does to the other reports - **Strategy comparison**: now that `recallAt5` has headroom it becomes the deciding metric again, and **hybrid does not clear the rule** — it matches dense on Recall@5 (0.9211) while improving nDCG@10 (0.8574 vs 0.8476), MRR and MAP@10. So the amended rule gives the conservative answer in the unsaturated regime and the permissive one in the saturated regime; both are reported rather than one being quietly preferred. - **Sweep**: `candidateK` finally moves something — dense nDCG@10 0.8212 at `candidateK=5` versus 0.8476 at 10 and above. `contextK=8` now buys context recall (1.0000 versus 0.9211) at a precision cost (0.3246 → 0.1316), which is the trade the column exists for. - **Threshold**: still flat from 0 to 0.6. Even an out-of-scope question scores above 0.6 against every chunk, so the threshold remains unmeasurable on a 19-chunk corpus. ## Testing - `npm test` — 504 pass - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` - `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated - `eval/splits.json` extended: 3 of the 8 new questions on `validation`, 5 on `test` ## Scope Hard negatives and near-duplicate *documents* are still not added; this increment is about question coverage over the existing corpus. The corpus is still 19 chunks, which is why `threshold` cannot be measured and `candidateK` saturates at 10. Part of #192 (child 2, stage 2: cross-lingual coverage).
`hybrid` runs a dense leg, and that leg applies a similarity floor. The trace recorded `threshold: undefined` for hybrid, so a snapshot could not reproduce the retrieval it described: it claimed hybrid had no threshold while `candidateHits()` was quietly applying `request.threshold ?? 0.5`. If hybrid ever becomes the shipped strategy, the #157 promise that a retrieval snapshot is reproducible stops holding. Found in the #192 review. Two separate `?? 0.5` defaults were the cause — one inside `candidateHits()`, one implicit in what the trace wrote. There is now **one** function, `denseChannelThreshold(strategy, requested)`, and the same value is both handed to the dense channel and written to the trace, so the two cannot drift apart again: - `dense` / `hybrid` → the requested threshold, or `DEFAULT_DENSE_THRESHOLD` - `sparse` → `undefined` (BM25 has no similarity floor) The trace field is renamed `threshold` → **`denseThreshold`**, because a bare `threshold` on a `strategy: 'hybrid'` trace reads as "the threshold for all of hybrid", which is the ambiguity that hid this bug. `parseRetrievalSnapshot` reads the legacy `threshold` from already-persisted snapshots as a dense threshold — the value was always the dense leg's, so the backfill states what that retrieval actually did. ## Testing - `npm run typecheck` — clean - `npm test` — 502 pass, with new coverage for the three cases above and a regression test asserting `hybrid` traces `0.5` while `sparse` traces nothing - Verified end to end that `dense` still reports `denseThreshold` in the smoke path ## Not covered `HybridRetriever` needs a vector store, so it has no unit test here; the pure decision is pinned instead and the wiring uses a single variable, which is what makes the divergence impossible rather than merely tested against. Part of #192 (review follow-up).
…ates "passages" Four follow-ups from the #192 review. The first two are semantics that would have spread into corpus v2 if they were left alone. **Strategy adoption is now held out.** `eval-retrieval.mjs` ran with `split = all`, so the strategy was chosen and scored on the same 44 questions — the exact mistake the threshold experiment had already been fixed for. It now: 1. runs every strategy on **validation** and decides there with `decideAdoption`; 2. re-runs the shipped strategy and the selected one on **test**, and only reports them. The choice never sees `test`. If validation selects nothing, `test` reports the shipped strategy alone and the report says there is no adoption candidate. The consequence is a more conservative and more trustworthy result than before: on validation, hybrid is **identical** to dense (Recall@5 0.9231, nDCG@10 0.8276 on both), so nothing is adopted — whereas the `split = all` run had hybrid ahead on nDCG@10. The old number was the choice being scored on its own questions. **`meanRetrieved` was a misleading name.** The harness fetches `candidateK` in order to compute `Recall@10`; the context window is `results.slice(0, contextK)`. So `meanRetrieved = 19` never meant "19 passages go to the model" — it meant "19 candidates passed the threshold". The unanswerable group now reports both sizes: ``` meanCandidatesRetrieved 19 // passed the threshold, capped by candidateK meanContextPassages 3 // actually reach the window, min(candidates, contextK) ``` The accurate description of the FIFA case is therefore: **19/19 chunks pass `threshold: 0.5`, and the top 3 irrelevant ones go into the prompt.** Still a real problem, but not "19 passages are stuffed into the model". **And it is retrieval abstention, not refusal.** No generator runs in this harness, so it can show that nothing passed the threshold; it cannot show that the model would decline to answer. Fields renamed (`abstentionCount`, `retrievalAbstentionRate`) and the report says plainly that a true system refusal rate needs a generator eval. **The manifest rationale was wrong.** The docs claimed a hash of the id "moves other questions between the sides" when one is added. That is false: `hash(id) % 3` is computed per id, so it is stable. The real reasons for an explicit manifest are that a hash cannot stratify a small corpus (which is how `multi-hop` and `cross-lingual` ended up entirely on one side) and that a new question would be assigned silently rather than deliberately. Corrected in `types.ts`, `eval/README.md` and the split tests. Regenerated: baseline, retrieval, threshold and sweep reports. The threshold table now shows `Unans. cands` and `Unans. ctx` as separate columns instead of conflating them. ## Testing - `npm run typecheck` — clean - `npm test` — 504 pass - `npm run eval` twice — byte-identical baseline - `npm run eval:retrieval` — validation selects, test reports; no adoption candidate - `npm run eval:threshold`, `npm run eval:sweep` — regenerated Part of #192 (review follow-ups).
…egatives
The first of the confusion clusters corpus v2 is built from. Not "near-duplicate documents
that lower Recall@5" — deliberately confusable material, so the benchmark can tell retrieval
strategies apart at all.
## The cluster
Four documents that share almost all their vocabulary and differ in every number:
| | v1 | v2 | v3 |
| --- | --- | --- | --- |
| Path | `/v1/complete` | `/v2/generate` | `/v3/chat` |
| Default timeout | 30000 ms | 60000 ms | 45000 ms |
| Recommended retries | 2 | 5 | 3 |
| Backoff base | 500 ms | 2000 ms | 1000 ms |
| Rate limit | 600 rpm | 3000 rpm | 1200 rpm |
| Context window | 8192 tokens | 32768 tokens | 65536 tokens |
All four documents discuss *timeout, retry, backoff, rate limit and context window* in that
order, so a query naming one field matches four passages and only one is right. A migration
guide restates every number in comparison tables, which is the hardest negative of the set:
it contains v1's timeout **and** v2's **and** v3's in one block.
## The tool that made it possible
`npm run eval:blocks <file.md>` prints the real `document_blocks.order` values through the
same loader and block builder the harness uses. Ground truth `block` is a pipeline ordinal,
not a line number a human counted; authoring a corpus by guessing it is how a dataset
quietly drifts. `eval/README.md` now says so and points at the command.
## What it measured
```
before after
chunks 19 35
Recall@5 0.9211 0.8913
nDCG@10 0.8476 0.7915
MAP@10 0.7965 0.7361
```
By type, the new cluster behaves as designed:
| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| hard-negative | 4 | **1.0000** | **0.7827** |
| multi-hop | 4 | 0.7500 | 0.7289 |
| semantic | 19 | 0.9474 | 0.8974 |
| cross-lingual | 9 | 0.6667 | 0.4311 |
`hard-negative` has perfect Recall@5 and a poor nDCG, which is the exact signature of this
kind of material: the correct passage **is** in the top 5, but a confusable sibling outranks
it. That is a ranking problem the strategy comparison can now see, and it is invisible on a
corpus where the right answer is always at rank 1.
`multi-hop` fell to 0.75 because the two comparison questions need four locations each and
do not get all four.
Also added: 8 questions in total (2 exact, 4 hard-negative, 1 semantic, 2 multi-hop, and two
**near-miss unanswerable** — "How much GPU memory does Gateway API v2 require?" and "What is
the monthly subscription price of Gateway API v3?", both about documents that discuss the
subject at length and never mention the answer). The FIFA questions were a sanity check for
a totally out-of-scope query; these are the ones that actually test abstention, because the
subject *is* in the library.
Unanswerable abstention is still 0/8, with `meanCandidatesRetrieved` at the `candidateK`
cap of 20 and `meanContextPassages` at 3.
## Scope, honestly
**35 chunks, not 150–300.** One cluster is not the target. The mechanism is proven and
verified end to end, and the remaining clusters are the same exercise repeated — but this PR
does not pretend the index is large enough for `candidateK ∈ {5,10,20,40}` to be a real
sweep. That arrives with the remaining clusters.
Part of #192 (child 2, stage 3: first confusion cluster).
Four more deliberately confusable documents, this time in the monitoring family the corpus
already had two members of, so the cluster is confusable with the *existing* documents as
well as internally.
| | reservoir | coastal | estuary | groundwater |
| --- | --- | --- | --- | --- |
| Sampling interval | fortnightly | hourly | daily | monthly |
| Replicates | 4 | 3 | 5 | 2 |
| Sensor depth | 5 m below surface | 1 m below surface | 2 m below surface | 15 m below water table |
Every document discusses *sampling interval, replicate samples and sensor depth* in the same
order, and each states the network's lowest or highest value ("the least frequent cadence in
the network", "the highest of any programme"), so a query lands on several plausible passages
and only one is right.
Corpus: 19 → 53 chunks, 13 → 21 documents, 54 → 63 questions.
## What it measured
```
after cluster 1 after cluster 2
chunks 35 53
Recall@5 0.8913 0.8774
nDCG@10 0.7915 0.7804
MAP@10 0.7361 0.7280
```
| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| cross-lingual | 9 | 0.6667 | 0.4173 |
| multi-hop | 5 | 0.7000 | 0.7218 |
| semantic | 21 | 0.9048 | 0.8542 |
| exact | 8 | 1.0000 | 0.8663 |
| hard-negative | 7 | **1.0000** | **0.8758** |
| zh | 3 | 1.0000 | 1.0000 |
`hard-negative` keeps the signature the cluster was built for: the correct passage is always
in the top 5 and a confusable sibling keeps outranking it.
## A result worth pausing on
The sweep (all 63 questions) now shows **hybrid ahead of dense on the deciding metric**:
| | Recall@5 | nDCG@10 |
| --- | --- | --- |
| dense, candidateK=20 | 0.8774 | 0.7804 |
| hybrid, candidateK=20 | **0.8962** | **0.8098** |
| hybrid, candidateK=5 | **0.9104** | 0.8101 |
But on the **validation** split the same comparison is a dead heat (both 0.8684), so the
held-out rule still declines to adopt.
That discrepancy is the honest state of things, and it is a property of the split rather
than of hybrid: 24 validation questions is not enough to see a gain that the full set shows.
The conclusion is to grow the validation split, **not** to go back to selecting on
everything — which is what produced the earlier, over-confident "hybrid clears the rule".
## Unanswerable is still 0/10
`meanCandidatesRetrieved` sits at the `candidateK` cap and `meanContextPassages` at
`contextK` for every unanswerable question, so abstention remains unmeasurable here. Two of
the ten are near-miss questions ("How much GPU memory does Gateway API v2 require?", "How
many litres per second does the reservoir release downstream?") about subjects the corpus
discusses at length — the realistic hallucination shape rather than the FIFA sanity check.
## Scope, honestly
**53 chunks, not 150–300.** Two clusters is not the target either. `candidateK ∈
{5,10,20,40}` now has more room than it did at 19 chunks, but `candidateK ≥ 10` still
behaves identically, so the axis is only partly unlocked. The remaining clusters are the same
exercise and the same verification loop; this PR stops where the verified work stops.
Part of #192 (child 2, stage 3: confusion clusters).
Adds `npm run eval:scores`, which measures whether a single dense similarity threshold is capable of separating "this passage answers the question" from "this one does not" — before any threshold grid is chosen. Child 2 of #192, and it changes what the next step should be. ## First, the score is not a cosine `SQLiteVectorStore` computes `score = 1 - distance / 2`, and sqlite-vec's cosine distance is `1 - cosine`, so: ``` score = (1 + cosine) / 2 => threshold 0.5 == cosine 0.0 ``` The shipped `threshold = 0.5` is therefore **cosine ≥ 0**, not "cosine ≥ 0.5". Every guard that mattered here — the sweep grid, the threshold report, the source comment — was describing an affine map, not the cosine. Fixed in the store's comment, in `eval/README.md`, and reported as two columns throughout the new diagnostics. ## The measurement Dense only (hybrid's `score` is an RRF value, `1 / (60 + rank)`, and not on the same scale), `validation` split only (the distribution is used to choose where to sweep, so it must not see `test`), `threshold = 0` and `candidateK = 500` (above the 53-chunk index, so every chunk is scored for every query). ``` best relevant p10 0.8906 (cosine 0.781) p50 0.9359 (0.872) best non-relevant p50 0.9243 (cosine 0.849) p90 0.9386 (0.877) margin p10 -0.0272 p50 +0.0026 unanswerable max p50 0.9259 (cosine 0.852) p90 0.9424 (0.885) ``` Three things fall out of that: 1. **The best non-relevant passage outranks the best relevant one about half the time.** The margin's p50 is +0.0026 and its p10 is −0.0272. Dense similarity is barely informative about relevance on this corpus — the hard-negative clusters are doing exactly what they were built for. 2. **The unanswerable max sits above the worst required relevant passage** (p90 0.9424 vs p10 0.8906), so the two distributions are not separable. 3. **`cross-lingual` is the lowest of every type** (best relevant p10 0.8808) against `semantic` at 0.9208 — a threshold tuned on the aggregate would cut cross-lingual first. ## The threshold curve is the answer | threshold | raw cosine | answerable hit | answerable full recall | unanswerable abstain | | --- | --- | --- | --- | --- | | 0.875 | 0.75 | 1.0000 | 1.0000 | 0.0000 | | 0.900 | 0.80 | 0.8947 | 0.8947 | 0.2000 | | 0.925 | 0.85 | 0.6316 | 0.6316 | 0.4000 | | 0.950 | 0.900 | 0.1579 | 0.1579 | 1.0000 | **Abstention never rises without full recall falling.** Every threshold that refuses an unanswerable question refuses required relevant passages at the same rate. There is no operating point that buys the first without paying the second. So the conclusion is not "sweep 0.7–0.8 instead of 0–0.6". It is: **a single dense similarity threshold cannot carry both recall and abstention on this corpus**, and the next mechanism to evaluate is a different signal — reranker score, top1−top2 margin, per-query thresholds, or claim-level answerability. That is a much more useful finding than a best threshold would have been, and it is why this PR was worth doing before more clusters. ## What this PR does not do - It does not choose a threshold. That is now blocked on a signal that can separate the distributions. - The `eval:threshold` sweep grid is left as it is. Re-gridding it would move a knob whose curve is flat where it matters and precipitous where it does not — the diagnostic is the replacement for re-gridding, not a preamble to it. ## Testing - `npm run typecheck` — clean - `npm test` — 504 pass - `npm run eval:scores` — writes `docs/eval/scores-v1.6.{json,md}` - The committed baseline is **unchanged**: `--eval-scores` is off by default, so the CI-diffed JSON does not grow a score series it has no use for Part of #192 (child 2: corpus diagnostics).
…ne transform
`score = (1 + cosine) / 2` is an affine map, so an **absolute** score converts as
`cosine = 2·score − 1`. A **difference** of scores does not: the `+1` cancels and
`Δcosine = 2·Δscore`.
The margin row applied the absolute transform, which reported `2Δscore − 1` — for a
margin near zero that is a cosine margin near **−1**, a sign flip on top of a scale error:
```
before after
margin p10 -0.0272 (-1.054) -0.0272 (-0.054)
margin p50 +0.0026 (-0.995) +0.0026 (+0.005)
```
`quantileRow` now takes the cosine transform as an argument and the margin row passes
`toCosineMargin`, so the two cannot be confused at the call site again. The report also
says which mapping applies to which row.
This does not change the separability conclusion — the p90 for the best non-relevant
passage is 0.9386 against a worst-required-relevant p10 of 0.8906, and that comparison
uses absolute scores, which were always converted correctly.
The margin itself is also now labelled in the report as an **oracle** quantity: at runtime
nothing knows which result is relevant, so it measures how much the score separates the
two, and it is not a signal the product could use.
Part of #192 (correction to the score diagnostics).
…stions `npm run eval:paired` reports the dense ↔ hybrid pair per question and the classification per query type. It answers the one thing the aggregate comparison could not. The open question was: hybrid is ahead on the full set and level on `validation`, and an average cannot say whether that is a broad small gain or a handful of rescued cases. Those two readings imply different next steps. Child 2 of #192. ## What it found ``` validation 19 questions improved 2 / tied 16 / regressed 1 mean ΔnDCG@10 +0.0047 test 34 questions improved 5 / tied 29 / regressed 0 mean ΔnDCG@10 +0.0432 ``` **The gain is not broad — it is five questions on `test` and two on `validation`.** On the reporting side, 29 of 34 questions are untouched. The mean ΔnDCG@10 of +0.0432 is carried by `q013` (rank 11 → 4, ΔnDCG +0.43), `q059` (5 → 2), `q026` (2 → 1), `q049` (3 → 2) and `q052` (2 → 1). ## And it is not the story we had started telling ourselves The plausible story was "hybrid helps the hard negatives". The breakdown says otherwise: | type | n (test) | improved | tied | regressed | | --- | --- | --- | --- | --- | | cross-lingual | 5 | **0** | 5 | 0 | | hard-negative | 5 | **1** | 4 | 0 | | multi-hop | 3 | 1 | 2 | 0 | | semantic | 15 | **3** | 12 | 0 | | exact | 4 | 0 | 4 | 0 | Three of the five wins are `semantic`, one is `hard-negative` and one `multi-hop`. With n=1 of 5, "hybrid helps hard-negative questions" is **not** supported at this size — which is exactly the pattern that would have been asserted from the aggregate number alone. **`cross-lingual` is untouched on both splits: 0 improved, 0 regressed.** Adding a sparse channel does nothing for the category that is measurably weakest, which is itself a useful negative result and a reason not to reach for hybrid as the answer to it. No question was found by one strategy and missed by the other, on either split, so no regression hides behind the `0 = not found` sentinel. ## How Metrics come from `src/main/eval/metrics.ts` — imported, not reimplemented — so a delta cannot disagree with the metric it is a delta of. `firstRelevantRank` is `0` for "never found", so the miss cases are classified explicitly rather than folded into an arithmetic delta that cannot express them. `validation` is the selecting side; the `test` breakdown **explains** the observed difference and is labelled as not-for-selection in the report itself. Nothing here changes the shipped strategy. ## Testing - `npm run typecheck` — clean - `npm run eval:paired` — 4 harness runs, writes `docs/eval/paired-v1.6.{json,md}` - The committed baseline is unchanged Part of #192 (child 2: corpus diagnostics).
…al PR Closes RAG Eval v2 Phase 1. This is the final expansion of the eval system — the conclusions below are the point of the phase, not more infrastructure. ## What was added Fifteen unanswerable questions whose **subject is discussed at length in the corpus and whose answer is not there** — "How much GPU memory does Gateway API v3 need?", "Which laboratory is accredited to analyse the estuary transect samples?", "How many staff work on the reservoir monitoring programme?". Every specific term was checked absent before the question was written (`gpu`, `vram`, `sla`, `uptime`, `expiry`, `sdk`, `self-hosted`, `accredit`, `vendor`, `supplier`, `funding`, `grant`, `staff`, `warranty`, `retention`, `encrypt`, `pricing`, `invoice`, `audit` all return nothing). That check is what makes them unanswerable rather than believed to be, and it is the part that cannot be automated away. `validation` unanswerable goes from **5 to 15**, so abstention resolution goes from 20 % per question to 6.7 %. The earlier "Who won the 2018 FIFA World Cup?" questions were a sanity check for a query that is obviously out of scope; these are the realistic shape — the user's actual hallucination risk is asking a question about a document that *is* in the library. ## The conclusion survives the larger sample ``` unanswerable max candidate min 0.8978 (cosine 0.796) p50 0.9279 (0.856) max 0.9435 (0.887) worst required relevant p10 0.8906 (cosine 0.781) ``` Even the **lowest** unanswerable top score is above the p10 of the worst required relevant passage. With n=15 instead of n=5 the distributions still do not separate, and the threshold curve is unchanged in shape: | threshold | raw cosine | answerable full recall | unanswerable abstain | | --- | --- | --- | --- | | 0.875 | 0.750 | 1.0000 | 0.0000 | | 0.900 | 0.800 | 0.8947 | 0.0667 | | 0.925 | 0.850 | 0.6316 | 0.4000 | | 0.950 | 0.900 | 0.1579 | 1.0000 | Abstention still never rises without full recall falling. The finer resolution makes the finding *more* solid, not different. Regenerated to stay consistent: baseline, scores, threshold and sweep. `paired-v1.6` is unchanged, because adding unanswerable questions does not touch any answerable one. ## Where Phase 1 leaves the product 1. The original 19-chunk benchmark could not guide a product decision; the current one (53 chunks, 78 questions, per-type metrics, held-out split) can. 2. **Cross-lingual retrieval is the largest measured gap**: Recall@5 0.6667 and nDCG@10 0.4173 against 0.93–1.00 for same-language questions. Chinese over a Chinese source scores 1.0000, so it is the cross-language matching, not Chinese. 3. **Hybrid has ranking value and no held-out mandate.** Its full-set advantage is five questions out of 34 (`paired-v1.6`), mostly `semantic`, and it does nothing for cross-lingual. Dense stays the default. 4. `contextK = 3` is justified against 5 and 8 on this corpus; `candidateK` above 10 is not distinguishable. 5. `threshold = 0.5` means **cosine ≥ 0**, not cosine ≥ 0.5. 6. **A single embedding similarity threshold cannot decide answerability** — the relevant and non-relevant and unanswerable score distributions overlap, and every threshold that abstains also drops required evidence. (6) is the phase's real result: it says stop tuning this knob and reach for a different mechanism. ## Stop line Phase 1 ends here. Not done, and deliberately not started: public benchmarks, reranker eval, generator faithfulness, citation entailment, growing the corpus to 300 chunks, an `Abstention Signal Evaluation`. Those build a more complete benchmark; they do not unblock KnowNote development, and (6) already says the next useful work is a different retrieval mechanism rather than a better measurement of this one.
6 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Minimal closeout of #192: integrate the already-completed stacked retrieval eval work (#195–#208) into main. Those PRs merged into preceding feature branches, not main. This PR preserves #209–#212 product fixes and ends the Phase 1 retrieval-eval expansion.
Closes #192.
Changes
Validation
No production strategy, K or threshold change. No generator/citation/public/reranker completion claim.