Skip to content

chore(eval): integrate Retrieval Eval v2 and close Phase 1 - #215

Merged
mrsibe merged 19 commits into
mainfrom
chore/eval-v2-closeout
Oct 1, 2026
Merged

mrsibe merged 19 commits into
mainfrom
chore/eval-v2-closeout

Conversation

@mrsibe

@mrsibe mrsibe commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Purpose

Minimal closeout of #192: integrate the already-completed stacked retrieval eval work (#195–#208) into main. Those PRs merged into preceding feature branches, not main. This PR preserves #209–#212 product fixes and ends the Phase 1 retrieval-eval expansion.

Closes #192.

Changes

  • Integrate existing metrics, split/validation selection, saturation-aware adoption, corpus and diagnostic reports; no new tuning or experiments.
  • Regenerate retrieval-v1.6 for the current 34 validation / 44 test questions; no winner and production dense defaults remain unchanged.
  • Resolve merge conflicts and remove a duplicated answerSources test.
  • Fix existing undefined winner/production references in the threshold report's non-flat branch and remove an unused paired-report helper.
  • Add docs/eval/README.md with Phase 1 boundaries and explicit limitations (21 docs / 53 chunks / 78 questions; original ~300-chunk target unmet and deferred).
  • Generator evaluation moves to [Eval] Generator evaluation — groundedness and answer completeness #213; citation evaluation moves to [Eval] Citation evaluation — precision, recall and entailment #214. Public suites, reranker and further scaling remain deferred, not closure gates.

Validation

  • npm test: 518 pass, zero failures.
  • npm run typecheck: pass.
  • npm run lint: zero errors, 110 existing warnings.
  • npm run build: pass.
  • npm run eval:retrieval: regenerated current report.
  • npm run eval:threshold: pass; flat sweep and generated reports unchanged.
  • Worker verified baseline JSON byte-for-byte reproduction with npm run eval.
  • git diff --check / staged diff check: pass.

No production strategy, K or threshold change. No generator/citation/public/reranker completion claim.

Two gates the v1.5 eval path needs to be trustworthy on every platform.

**`npm run eval` could not start on Windows.** The three scripts resolve
`node_modules/.bin/electron.cmd` and spawn it. Node 24 refuses to spawn a
`.cmd`/`.bat` without `shell: true` and fails with `EINVAL`, so the harness was
unrunnable on the platform this project is developed on. The `electron` package
exports the path to the real executable the wrapper runs, which spawns directly
on every platform and needs no shell.

**A Windows run rewrote the committed baseline.** `path.relative` returns
backslashes on Windows, so `config.corpus` was written as `eval\corpus` instead
of `eval/corpus`. The file's whole contract is that it is identical on every
machine — the CI determinism check diffs it — and a Windows run silently broke
that. The label is now normalised to POSIX separators.

Verified on Windows with the pinned model: `npm run eval` runs and leaves
`docs/eval/baseline-v1.5.json` byte-identical to the committed file.
… config

The harness measured a retriever nobody runs. Production retrieved at `topK: 3`
with `threshold: 0.5`; the harness ran `topK: 10` with `threshold: 0` and reported
`Recall@5 = 1.0000` for a pipeline that silently drops rank 4 at cosine 0.47. The
first child of #192 asks for the two to be the same configuration.

**One K was doing two jobs.** `RetrievalRequest.topK` was both the first-stage
width (KNN neighbours, BM25 limit) and the number of passages delivered. In hybrid
that made fusion nearly a no-op: dense contributed `topK`, BM25 contributed
`topK`, RRF fused at most `2 * topK`, and the result was sliced straight back to
`topK`. The two stages are now named:

```
RetrievalRequest:
  candidateK  first-stage width per channel   (default 20)
  topK        passages delivered              (default 5)
```

`effectiveCandidateK` enforces `candidateK >= topK`, so a caller that asks for
more results than the default pool (the search palette, MCP) is never silently
capped. `HybridRetriever` fuses the wide pool and truncates once, after fusion.
`DenseRetriever` queries `candidateK` and slices to `topK` — equivalent to before
for a single strategy, since a prefix of a ranking is the same ranking.

The trace and the #157 snapshot gained `candidateK`. A snapshot written before
this change backfills it from `topK`, which is what that retrieval actually did.

**The harness now defaults to production** (`candidateK: 20`, `contextK: 3`,
`threshold: 0.5`), with `--eval-candidate-k=` / `--eval-context-k=` /
`--eval-threshold=` to move them deliberately. Ranking metrics are computed at
`candidateK` depth, not `contextK`: `Recall@10` needs ten results, and truncation
only takes a prefix, so it cannot change the ranking being measured. `contextK` is
recorded so the report describes the whole online path.

`docs/eval/baseline-v1.6.{json,md}` is the new frozen baseline; v1.5 is kept as
history for the #77/#78 deltas, and the CI determinism check moves to v1.6.

## What this did and did not change

On the current 19-chunk corpus the production configuration produces the **same**
metrics as v1.5 — `threshold: 0.5` is non-binding and `candidateK: 20` exceeds the
corpus, so nothing is filtered and no ranking changes. That is the corpus
limitation #192 describes, not a result. What changed is that the harness now
*states* the production parameters instead of assuming different ones.

Hybrid still leads dense on Recall@1 (0.8667 vs 0.8333), MRR (0.9444 vs 0.9278)
and nDCG@10 (0.9561 vs 0.9437) with a wide pool; dense stays the default because
the adoption rule still cannot move on a saturated Recall@5.

## Testing

- `npm run typecheck` — clean
- `npm test` — 481 pass (3 new: trace keeps the two Ks apart, the width invariant,
  the pre-#77 snapshot backfill)
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`
- `npm run eval:retrieval` — dense/sparse/hybrid unchanged in ordering

Part of #192 (child 1).
Recall@K, MRR and nDCG@10 could not express three things the epic needs, and a
single aggregate average was hiding a gap the corpus already contained. Child 3 of
#192.

**Two metrics added.** `hitRate@5` is whether *any* of the first five passages
covers ground truth; `MAP@10` combines ranking position with coverage, so pulling a
second relevant passage from rank 9 to rank 2 moves it while Recall@5 sits still.
Hit rate is deliberately blunt next to Recall@5: a two-passage question that finds
one scores 1.0 and 0.5 respectively, and both facts matter — "the model had a
chance" is not "the material was complete".

**Every metric is now reported per query type.** `questions.jsonl` gained an optional
`type`, the harness groups by it, and the report renders a table. The aggregate was
already concealing something:

| Type | n | Recall@5 | nDCG@10 | Hit rate@5 | MAP@10 |
| --- | --- | --- | --- | --- | --- |
| cross-lingual | 1 | 1.0000 | **0.6309** | 1.0000 | **0.5000** |
| exact | 6 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| multi-hop | 2 | 1.0000 | 0.9599 | 1.0000 | 0.9167 |
| semantic | 18 | 1.0000 | 0.9312 | 1.0000 | 0.9074 |
| zh | 3 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| **all** | 30 | 1.0000 | 0.9437 | 1.0000 | 0.9222 |

The one cross-lingual question (`q030`, Chinese over an English source) ranks far
worse than everything else. Recall@5 = 1.0000 reported that as a success; nDCG@10
and MAP@10 are what make the multilingual gap visible. That is the metric doing its
job on the existing 30 questions, before the corpus grows.

Untagged questions group under `untagged` rather than being dropped, and the types
in use are documented in `eval/README.md`.

## Testing

- `npm run typecheck` — clean
- `npm test` — 485 pass, 4 new: hit rate vs recall, hit rate@k bounds, AP position
  sensitivity, AP's repeat-counts-once rule
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`

Part of #192 (child 3).
…test

`threshold: 0.5` was hand-picked, and a cosine score has no universal meaning — its
distribution depends on the embedding model, the language, the query type and the
chunk length. Child 4 of #192.

**A deterministic split, owned by the harness.** `--eval-split=validation|test`
partitions the questions by a hash of the id, so the same `questions.jsonl` cuts the
same way on every machine and both arms go through the same code path that produces
the frozen baseline. A parameter chosen on the questions it is scored on is fitted,
not measured.

**`npm run eval:threshold`** sweeps `0 / 0.3 / 0.4 / 0.5 / 0.6`, selects on
`validation`, and reports the winner on `test`. It also counts a metric the baseline
does not carry: the **no-result rate**. A higher threshold can look better on a
ranking metric while quietly making the product answer "not in your sources" more
often, and that trade is invisible unless it is counted.

## The result here is a non-result, and it is reported as one

| Threshold | Recall@5 (val) | nDCG@10 (val) | No-result (val) | nDCG@10 (test) | No-result (test) |
| --- | --- | --- | --- | --- | --- |
| 0 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.3 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.4 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.5 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.6 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |

The sweep is **flat**: every threshold produces identical metrics and never filters a
passage. E5 does not score these query/chunk pairs below 0.6, so on this corpus the
threshold is non-binding. The report says **"no evidence to change
`threshold = 0.5`"** rather than nominating the tie-break winner, because moving a
product parameter on a flat sweep would be noise dressed as a result. The tie-break
rule (widest threshold among equals) is stated so it can be argued with.

That is the honest outcome, and it is another instance of the corpus limitation
#192 describes. The machinery is what ships here; the number is what the corpus
cannot yet support.

## Testing

- `npm run typecheck` — clean
- `npm test` — 489 pass, 4 new: split bounds, split determinism, validation/test
  partition without overlap, `all` preserves order
- `npm run eval` — `docs/eval/baseline-v1.6.json` regenerated with `config.split`
- `npm run eval:threshold` — reruns cleanly and rewrites the same tables

Part of #192 (child 4).
The v1.5 rule was "Recall@5 must improve and nDCG@10 must not regress". On a corpus
where dense already scores Recall@5 = 1.0000, no strategy can improve Recall@5, so
the rule was not strict — it was **unsatisfiable**. Every comparison came back
"inconclusive", including a hybrid that was better on Recall@1, MRR and nDCG@10. The
default never moved, not because hybrid lost but because the rule could not return a
verdict. Child 10 of #192.

**The amendment.** A metric at its maximum has no headroom and is not allowed to
decide. The deciding metric is the first one with headroom, in the order
`recallAt5`, `ndcgAt10`, `mrr`, `mapAt10`; a strategy clears the rule when it improves
that metric and regresses none of the others. Saturation is detected and reported
rather than silently blocking every change.

**The rule is code, not prose.** It lives in `src/main/eval/adoption.ts` with 8 unit
tests, because it decides whether a shipped default moves and testing it by reading
the sentence the script prints would only test the sentence. The experiment script
imports it, so there is one definition.

## The v1.5 stalemate resolves

```
| Strategy                    | Recall@1 | Recall@5 | MRR    | nDCG@10 | MAP@10 | Query p95 |
| dense (vector)              | 0.8333   | 1.0000   | 0.9278 | 0.9437  | 0.9222 | 14.06 ms  |
| sparse (BM25)               | 0.7333   | 0.8667   | 0.8056 | 0.8184  | 0.8000 | 2.63 ms   |
| hybrid (RRF dense + BM25)   | 0.8667   | 1.0000   | 0.9444 | 0.9561  | 0.9389 | 22.32 ms  |
```

`recallAt5` is reported as saturated; the deciding metric is `ndcgAt10`; **hybrid
clears the rule** and regresses none of the four metrics, at a p95 cost of ~8 ms.

The script reports the measurement and does not flip the default — changing the
shipped strategy is a product decision, and it is stated that way in the output
rather than implied by a green checkmark.

`docs/eval/retrieval-v1.6.{json,md}` is the regenerated experiment against the v1.6
baseline; `retrieval-v1.5.*` is kept as the record of the superseded rule.

## Testing

- `npm run typecheck` — clean
- `npm test` — 497 pass, 8 new: saturation detection, deciding-metric priority, the
  float-slack regression check, the resolved stalemate, a trade that regresses another
  metric being refused, the baseline not clearing against itself, all-saturated, and
  best-candidate selection
- `npm run eval` + `npm run eval:retrieval` — regenerated and re-run

Part of #192 (child 10).
…odel

The harness measured the retriever but never the window the prompt actually gets.
`evidenceK: 5` was a separate constant from the `contextK: 3` production uses, so the
one metric that looked at a window looked at a different one than the product does.
Child 5 of #192.

**`contextK` is now the window for both context metrics.**

- `contextPrecision@contextK` — of the first `contextK` passages, the share covering
  ground truth. This is what `evidencePrecisionAt5` was, at the production width.
- `contextRecall@contextK` — the share of needed ground-truth blocks that made it into
  that window. Distinct from `Recall@10`: a block found at rank 4 is invisible when
  `contextK = 3`, and that is a product fact, not a ranking fact.

Both are **deterministic**: the dataset says which blocks answer the question, so no
model is needed to score a window. Together they are the trade-off a `contextK`
decision makes — a wider window finds more and carries more noise — which is what the
sweep in the next child needs.

Baseline: `contextPrecision@3 = 0.3556`, `contextRecall@3 = 1.0000` — the needed
evidence is always inside the top 3 on this corpus, but only about a third of what is
inside the window is relevant. That second number is the one that says the window is
paying for passages that do not answer the question.

**What is deliberately not here.** Faithfulness, completeness, answer correctness and
noise sensitivity need a generative model. The harness runs offline with only the
pinned embedding model — the same constraint that keeps the reranker unmeasured
(#170) — so this PR does not add them and does not fake them: the report's
Definitions section now says so, and the "Not evaluated" line in
`eval-retrieval.mjs` points at the same constraint. Adding an LLM judge is a separate
change that has to solve model pinning first, not a line of code.

`evidenceK` is removed from the harness config, so there is exactly one context width.

## Testing

- `npm run typecheck` — clean
- `npm test` — 497 pass
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`
- `npm run eval:retrieval` — regenerated; hybrid still clears the amended rule

Part of #192 (child 5, deterministic half).
The epic asks for a grid over `candidateK` and `contextK` with latency, index size and
context size recorded next to quality, and explicitly not for a single aggregate
"RAG score". Child 7 of #192.

`npm run eval:sweep` runs the real harness over
`strategy × candidateK {5,10,20,40} × contextK {3,5,8}` (24 runs, chunking fixed) and
writes `docs/eval/sweep-v1.6.{json,md}`. The harness gained `contextChars` per
question — the size of the context window, reported as **characters, not tokens**,
because the harness pins an embedding model and no generation tokenizer.

## What the dashboard already shows

Two of the three axes are decided by this corpus, and one of them decisively:

- **`candidateK` changes nothing.** Every metric is identical from 5 to 40, because
  the corpus is 19 chunks and the relevant passages are already inside the top 5. This
  is the saturation problem #192 describes, now visible on the axis it affects.
- **`contextK` has a clear optimum here: 3.** Context recall is 1.0000 at every width,
  while precision falls `0.3556 → 0.2133 → 0.1333` and the window grows
  `2079 → 3327 → 5312` characters as it widens. Wider adds prompt cost and noise for
  **no** recall. The production default is already 3, and this is the first evidence
  that it is the right 3 rather than a guess.
- **hybrid beats dense** on nDCG@10 (0.9561 vs 0.9437) and MAP@10 (0.9389 vs 0.9222) at
  every setting, with no metric regressing — consistent with #197.

The report labels the grid maximum as **not** a recommendation, because selecting on the
same questions is how a benchmark becomes a lookup table; the adoption rule and the
`validation`/`test` split are what keep the decision honest.

Timing p95 is reported per row but is noisy at this sample size; it is in the table so
a latency cost can be seen, not so it can be ranked.

## Testing

- `npm run typecheck` — clean
- `npm test` — 497 pass
- `npm run eval` — regenerated (`contextChars` added to `perQuestion`)
- `npm run eval:sweep` — 24 runs, dashboard written

Part of #192 (child 7).
The sweep contained cells the harness cannot fill. The harness retrieves `candidateK`
passages and the context metrics look at the first `contextK` of them, so
`candidateK=5, contextK=8` reports on five passages while claiming eight. The dashboard
showed `contextK=5` and `contextK=8` at `candidateK=5` as **identical** — and because
`evidencePrecisionAtK` divides by the passages actually retrieved, nothing exposed it.

Found in the #192 review. A wrong number that looks like a measurement is worse than a
failure, so the invariant is enforced rather than documented:

- **The harness refuses it.** `contextK > candidateK` throws before indexing, naming
  both numbers:
  `contextK (8) cannot exceed candidateK (5): the harness retrieves candidateK
  passages, so a wider context window can never be filled.`
- **The sweep skips those cells** and says so. The report lists the skipped
  combinations in a "Skipped cells" section instead of quietly omitting rows, because
  "we did not measure this" and "this measured the same as its neighbour" are different
  statements and the old output showed the second while meaning the first.

`docs/eval/sweep-v1.6.*` is regenerated: 22 rows instead of 24, and the misleading
`candidateK=5 / contextK=8` row is gone rather than silently equal to `5`.

The production cells (`candidateK=20, contextK ∈ {3,5,8}`) are unaffected.

## Testing

- `npm run typecheck` — clean
- `npm test` — 502 pass, including 3 new harness-invariant tests that need no vector
  store, because the check runs before any database work: the refusal, the message
  naming both numbers, and `contextK == candidateK` being accepted
- `node_modules/electron/dist/electron.exe . --eval-harness --eval-candidate-k=5
  --eval-context-k=8` — refused with the message above
- `npm run eval:sweep` — 22 rows, 2 skipped and listed

Part of #192 (review follow-up).
The corpus contract corpus v2 needs. Without it, two of the four things #192 asks the
dataset to cover cannot be represented at all: a question whose right answer is "your
sources do not say", and a validation/test assignment that is chosen rather than
computed.

**Unanswerable questions.** `answerable: false` with `relevant: []`. They are excluded
from `metrics` and `byType` — `recallAtK` treats no ground truth as 0/0, and averaging
"correctly refused" into "missed" would corrupt the table — and reported as their own
group: `noResultCount`, `noResultRate` and `meanRetrieved`. On this row **higher
no-result is better**, which is the opposite of how it reads everywhere else.
`assertQuestionShape` refuses the reverse combination in either direction, because both
are silent: an answerable question with no ground truth reads as a permanent miss, and
an unanswerable one carrying ground truth reads as a normal hit.

**An explicit split manifest.** `eval/splits.json` replaces the hash of the question id.
A hash is reproducible but not *stable*: adding a question moved others between the
sides, and a rare query type could end up entirely on one side without anyone choosing
that — which is what had happened (`multi-hop` and `cross-lingual` were both entirely on
`test`). A question with no manifest entry is now **refused** rather than defaulted, so a
new question cannot leak into the reporting side.

**Six unanswerable questions** are added (three topic-adjacent, three plainly
out-of-scope). Their specifics were checked absent from the corpus before being written,
so they are genuinely unanswerable rather than believed to be.

## What the first measurement shows

```
unanswerable: { questions: 6, noResultRate: 0, meanRetrieved: 19 }
```

Every unanswerable question returns **the entire corpus** — 19 of 19 chunks — at
`threshold: 0.5`, and the threshold sweep is flat all the way to `0.6`: refusal stays
`0/6` and `meanRetrieved` stays at the cap. So on this corpus the threshold cannot
separate a relevant passage from an unrelated one, and "Who won the 2018 FIFA World Cup?"
is answered with 19 passages of river monitoring, tidal energy and tea storage.

That is the evidence the threshold decision did not have before, and it is also the
corpus limitation #192 describes: it does not say `0.5` is wrong, it says this corpus
cannot tell. The reports say so rather than recommending a change.

Retrieval metrics are unchanged (`recallAt5` 1.0000), because the six new questions are
correctly kept out of them.

## Testing

- `npm run typecheck` — clean
- `npm test` — 504 pass, including the manifest contract (unknown side rejected, missing
  entry refused, `all` needs no manifest), the answerable/unanswerable tie, and the
  harness's window invariant
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`
- `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated; the
  threshold rule now gates on answerable quality and then maximises unanswerable refusal

## Scope

This is the **contract**, plus the smallest honest corpus increment that exercises it.
Hard negatives and near-duplicate documents — the thing that would bring `Recall@5` off
1.0000 and give `candidateK` something to do — are the next step, not this PR.

Part of #192 (child 2, stage 1 of the corpus).
…saturated

The corpus had exactly **one** cross-lingual question (`q030`, Chinese over an English
source), which is not enough to call anything a weakness — correctly flagged in the #192
review as a signal rather than a finding. This adds eight more Chinese questions over
**English** sources, reusing the already-verified ground truth of their English
counterparts: the fact is the same, only the query language changes, which is precisely
what the multilingual model has to bridge. Cross-lingual goes from n=1 to n=9.

## Recall@5 is no longer 1.0000

```
                 before    after
Recall@5         1.0000    0.9211
nDCG@10          0.9437    0.8476
MAP@10           0.9222    0.7965
hitRate@5        1.0000    0.9211
```

The benchmark was saturated on this corpus, which is why no retrieval change could ever
clear the adoption rule and why `candidateK` appeared to do nothing. Eight questions were
enough to bring Recall@5 off the ceiling.

## The finding, at a size that supports it

| type | n | Recall@5 | nDCG@10 | MAP@10 |
| --- | --- | --- | --- | --- |
| **cross-lingual** | **9** | **0.6667** | **0.5032** | **0.3446** |
| exact | 6 | 1.0000 | 1.0000 | 1.0000 |
| multi-hop | 2 | 1.0000 | 0.9599 | 0.9167 |
| semantic | 18 | 1.0000 | 0.9312 | 0.9074 |
| zh | 3 | 1.0000 | 1.0000 | 1.0000 |

Chinese questions over a **Chinese** source (`zh`) still score 1.0000. Chinese questions
over an **English** source score 0.6667. So the gap is specifically **cross-lingual
retrieval**, not Chinese, and it is now measured on nine questions rather than asserted
from one.

This is the kind of finding that was not defensible before: `q030` alone gave nDCG 0.63
and the honest conclusion was "a signal worth testing". With n=9 the same direction holds
at the same magnitude, so it is a real weakness of the current pipeline on this corpus.

## What it does to the other reports

- **Strategy comparison**: now that `recallAt5` has headroom it becomes the deciding
  metric again, and **hybrid does not clear the rule** — it matches dense on Recall@5
  (0.9211) while improving nDCG@10 (0.8574 vs 0.8476), MRR and MAP@10. So the amended rule
  gives the conservative answer in the unsaturated regime and the permissive one in the
  saturated regime; both are reported rather than one being quietly preferred.
- **Sweep**: `candidateK` finally moves something — dense nDCG@10 0.8212 at `candidateK=5`
  versus 0.8476 at 10 and above. `contextK=8` now buys context recall (1.0000 versus
  0.9211) at a precision cost (0.3246 → 0.1316), which is the trade the column exists for.
- **Threshold**: still flat from 0 to 0.6. Even an out-of-scope question scores above 0.6
  against every chunk, so the threshold remains unmeasurable on a 19-chunk corpus.

## Testing

- `npm test` — 504 pass
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`
- `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated
- `eval/splits.json` extended: 3 of the 8 new questions on `validation`, 5 on `test`

## Scope

Hard negatives and near-duplicate *documents* are still not added; this increment is about
question coverage over the existing corpus. The corpus is still 19 chunks, which is why
`threshold` cannot be measured and `candidateK` saturates at 10.

Part of #192 (child 2, stage 2: cross-lingual coverage).
`hybrid` runs a dense leg, and that leg applies a similarity floor. The trace
recorded `threshold: undefined` for hybrid, so a snapshot could not reproduce the
retrieval it described: it claimed hybrid had no threshold while `candidateHits()`
was quietly applying `request.threshold ?? 0.5`. If hybrid ever becomes the shipped
strategy, the #157 promise that a retrieval snapshot is reproducible stops holding.

Found in the #192 review.

Two separate `?? 0.5` defaults were the cause — one inside `candidateHits()`, one
implicit in what the trace wrote. There is now **one** function,
`denseChannelThreshold(strategy, requested)`, and the same value is both handed to the
dense channel and written to the trace, so the two cannot drift apart again:

- `dense` / `hybrid` → the requested threshold, or `DEFAULT_DENSE_THRESHOLD`
- `sparse` → `undefined` (BM25 has no similarity floor)

The trace field is renamed `threshold` → **`denseThreshold`**, because a bare
`threshold` on a `strategy: 'hybrid'` trace reads as "the threshold for all of
hybrid", which is the ambiguity that hid this bug. `parseRetrievalSnapshot` reads the
legacy `threshold` from already-persisted snapshots as a dense threshold — the value
was always the dense leg's, so the backfill states what that retrieval actually did.

## Testing

- `npm run typecheck` — clean
- `npm test` — 502 pass, with new coverage for the three cases above and a regression
  test asserting `hybrid` traces `0.5` while `sparse` traces nothing
- Verified end to end that `dense` still reports `denseThreshold` in the smoke path

## Not covered

`HybridRetriever` needs a vector store, so it has no unit test here; the pure decision
is pinned instead and the wiring uses a single variable, which is what makes the
divergence impossible rather than merely tested against.

Part of #192 (review follow-up).
…ates "passages"

Four follow-ups from the #192 review. The first two are semantics that would have spread
into corpus v2 if they were left alone.

**Strategy adoption is now held out.** `eval-retrieval.mjs` ran with `split = all`, so the
strategy was chosen and scored on the same 44 questions — the exact mistake the threshold
experiment had already been fixed for. It now:

1. runs every strategy on **validation** and decides there with `decideAdoption`;
2. re-runs the shipped strategy and the selected one on **test**, and only reports them.

The choice never sees `test`. If validation selects nothing, `test` reports the shipped
strategy alone and the report says there is no adoption candidate.

The consequence is a more conservative and more trustworthy result than before: on
validation, hybrid is **identical** to dense (Recall@5 0.9231, nDCG@10 0.8276 on both), so
nothing is adopted — whereas the `split = all` run had hybrid ahead on nDCG@10. The old
number was the choice being scored on its own questions.

**`meanRetrieved` was a misleading name.** The harness fetches `candidateK` in order to
compute `Recall@10`; the context window is `results.slice(0, contextK)`. So
`meanRetrieved = 19` never meant "19 passages go to the model" — it meant "19 candidates
passed the threshold". The unanswerable group now reports both sizes:

```
meanCandidatesRetrieved  19   // passed the threshold, capped by candidateK
meanContextPassages       3   // actually reach the window, min(candidates, contextK)
```

The accurate description of the FIFA case is therefore: **19/19 chunks pass `threshold:
0.5`, and the top 3 irrelevant ones go into the prompt.** Still a real problem, but not
"19 passages are stuffed into the model".

**And it is retrieval abstention, not refusal.** No generator runs in this harness, so it
can show that nothing passed the threshold; it cannot show that the model would decline to
answer. Fields renamed (`abstentionCount`, `retrievalAbstentionRate`) and the report says
plainly that a true system refusal rate needs a generator eval.

**The manifest rationale was wrong.** The docs claimed a hash of the id "moves other
questions between the sides" when one is added. That is false: `hash(id) % 3` is computed
per id, so it is stable. The real reasons for an explicit manifest are that a hash cannot
stratify a small corpus (which is how `multi-hop` and `cross-lingual` ended up entirely on
one side) and that a new question would be assigned silently rather than deliberately.
Corrected in `types.ts`, `eval/README.md` and the split tests.

Regenerated: baseline, retrieval, threshold and sweep reports. The threshold table now
shows `Unans. cands` and `Unans. ctx` as separate columns instead of conflating them.

## Testing

- `npm run typecheck` — clean
- `npm test` — 504 pass
- `npm run eval` twice — byte-identical baseline
- `npm run eval:retrieval` — validation selects, test reports; no adoption candidate
- `npm run eval:threshold`, `npm run eval:sweep` — regenerated

Part of #192 (review follow-ups).
…egatives

The first of the confusion clusters corpus v2 is built from. Not "near-duplicate documents
that lower Recall@5" — deliberately confusable material, so the benchmark can tell retrieval
strategies apart at all.

## The cluster

Four documents that share almost all their vocabulary and differ in every number:

| | v1 | v2 | v3 |
| --- | --- | --- | --- |
| Path | `/v1/complete` | `/v2/generate` | `/v3/chat` |
| Default timeout | 30000 ms | 60000 ms | 45000 ms |
| Recommended retries | 2 | 5 | 3 |
| Backoff base | 500 ms | 2000 ms | 1000 ms |
| Rate limit | 600 rpm | 3000 rpm | 1200 rpm |
| Context window | 8192 tokens | 32768 tokens | 65536 tokens |

All four documents discuss *timeout, retry, backoff, rate limit and context window* in that
order, so a query naming one field matches four passages and only one is right. A migration
guide restates every number in comparison tables, which is the hardest negative of the set:
it contains v1's timeout **and** v2's **and** v3's in one block.

## The tool that made it possible

`npm run eval:blocks <file.md>` prints the real `document_blocks.order` values through the
same loader and block builder the harness uses. Ground truth `block` is a pipeline ordinal,
not a line number a human counted; authoring a corpus by guessing it is how a dataset
quietly drifts. `eval/README.md` now says so and points at the command.

## What it measured

```
                 before    after
chunks           19        35
Recall@5         0.9211    0.8913
nDCG@10          0.8476    0.7915
MAP@10           0.7965    0.7361
```

By type, the new cluster behaves as designed:

| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| hard-negative | 4 | **1.0000** | **0.7827** |
| multi-hop | 4 | 0.7500 | 0.7289 |
| semantic | 19 | 0.9474 | 0.8974 |
| cross-lingual | 9 | 0.6667 | 0.4311 |

`hard-negative` has perfect Recall@5 and a poor nDCG, which is the exact signature of this
kind of material: the correct passage **is** in the top 5, but a confusable sibling outranks
it. That is a ranking problem the strategy comparison can now see, and it is invisible on a
corpus where the right answer is always at rank 1.

`multi-hop` fell to 0.75 because the two comparison questions need four locations each and
do not get all four.

Also added: 8 questions in total (2 exact, 4 hard-negative, 1 semantic, 2 multi-hop, and two
**near-miss unanswerable** — "How much GPU memory does Gateway API v2 require?" and "What is
the monthly subscription price of Gateway API v3?", both about documents that discuss the
subject at length and never mention the answer). The FIFA questions were a sanity check for
a totally out-of-scope query; these are the ones that actually test abstention, because the
subject *is* in the library.

Unanswerable abstention is still 0/8, with `meanCandidatesRetrieved` at the `candidateK`
cap of 20 and `meanContextPassages` at 3.

## Scope, honestly

**35 chunks, not 150–300.** One cluster is not the target. The mechanism is proven and
verified end to end, and the remaining clusters are the same exercise repeated — but this PR
does not pretend the index is large enough for `candidateK ∈ {5,10,20,40}` to be a real
sweep. That arrives with the remaining clusters.

Part of #192 (child 2, stage 3: first confusion cluster).
Four more deliberately confusable documents, this time in the monitoring family the corpus
already had two members of, so the cluster is confusable with the *existing* documents as
well as internally.

| | reservoir | coastal | estuary | groundwater |
| --- | --- | --- | --- | --- |
| Sampling interval | fortnightly | hourly | daily | monthly |
| Replicates | 4 | 3 | 5 | 2 |
| Sensor depth | 5 m below surface | 1 m below surface | 2 m below surface | 15 m below water table |

Every document discusses *sampling interval, replicate samples and sensor depth* in the same
order, and each states the network's lowest or highest value ("the least frequent cadence in
the network", "the highest of any programme"), so a query lands on several plausible passages
and only one is right.

Corpus: 19 → 53 chunks, 13 → 21 documents, 54 → 63 questions.

## What it measured

```
                 after cluster 1   after cluster 2
chunks           35                53
Recall@5         0.8913            0.8774
nDCG@10          0.7915            0.7804
MAP@10           0.7361            0.7280
```

| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| cross-lingual | 9 | 0.6667 | 0.4173 |
| multi-hop | 5 | 0.7000 | 0.7218 |
| semantic | 21 | 0.9048 | 0.8542 |
| exact | 8 | 1.0000 | 0.8663 |
| hard-negative | 7 | **1.0000** | **0.8758** |
| zh | 3 | 1.0000 | 1.0000 |

`hard-negative` keeps the signature the cluster was built for: the correct passage is always
in the top 5 and a confusable sibling keeps outranking it.

## A result worth pausing on

The sweep (all 63 questions) now shows **hybrid ahead of dense on the deciding metric**:

| | Recall@5 | nDCG@10 |
| --- | --- | --- |
| dense, candidateK=20 | 0.8774 | 0.7804 |
| hybrid, candidateK=20 | **0.8962** | **0.8098** |
| hybrid, candidateK=5 | **0.9104** | 0.8101 |

But on the **validation** split the same comparison is a dead heat (both 0.8684), so the
held-out rule still declines to adopt.

That discrepancy is the honest state of things, and it is a property of the split rather
than of hybrid: 24 validation questions is not enough to see a gain that the full set shows.
The conclusion is to grow the validation split, **not** to go back to selecting on
everything — which is what produced the earlier, over-confident "hybrid clears the rule".

## Unanswerable is still 0/10

`meanCandidatesRetrieved` sits at the `candidateK` cap and `meanContextPassages` at
`contextK` for every unanswerable question, so abstention remains unmeasurable here. Two of
the ten are near-miss questions ("How much GPU memory does Gateway API v2 require?", "How
many litres per second does the reservoir release downstream?") about subjects the corpus
discusses at length — the realistic hallucination shape rather than the FIFA sanity check.

## Scope, honestly

**53 chunks, not 150–300.** Two clusters is not the target either. `candidateK ∈
{5,10,20,40}` now has more room than it did at 19 chunks, but `candidateK ≥ 10` still
behaves identically, so the axis is only partly unlocked. The remaining clusters are the same
exercise and the same verification loop; this PR stops where the verified work stops.

Part of #192 (child 2, stage 3: confusion clusters).
Adds `npm run eval:scores`, which measures whether a single dense similarity threshold is
capable of separating "this passage answers the question" from "this one does not" — before
any threshold grid is chosen. Child 2 of #192, and it changes what the next step should be.

## First, the score is not a cosine

`SQLiteVectorStore` computes `score = 1 - distance / 2`, and sqlite-vec's cosine distance is
`1 - cosine`, so:

```
score = (1 + cosine) / 2      =>      threshold 0.5 == cosine 0.0
```

The shipped `threshold = 0.5` is therefore **cosine ≥ 0**, not "cosine ≥ 0.5". Every guard
that mattered here — the sweep grid, the threshold report, the source comment — was
describing an affine map, not the cosine. Fixed in the store's comment, in `eval/README.md`,
and reported as two columns throughout the new diagnostics.

## The measurement

Dense only (hybrid's `score` is an RRF value, `1 / (60 + rank)`, and not on the same scale),
`validation` split only (the distribution is used to choose where to sweep, so it must not
see `test`), `threshold = 0` and `candidateK = 500` (above the 53-chunk index, so every chunk
is scored for every query).

```
best relevant     p10 0.8906 (cosine 0.781)   p50 0.9359 (0.872)
best non-relevant p50 0.9243 (cosine 0.849)   p90 0.9386 (0.877)
margin            p10 -0.0272                 p50 +0.0026
unanswerable max  p50 0.9259 (cosine 0.852)   p90 0.9424 (0.885)
```

Three things fall out of that:

1. **The best non-relevant passage outranks the best relevant one about half the time.**
   The margin's p50 is +0.0026 and its p10 is −0.0272. Dense similarity is barely
   informative about relevance on this corpus — the hard-negative clusters are doing exactly
   what they were built for.
2. **The unanswerable max sits above the worst required relevant passage** (p90 0.9424 vs
   p10 0.8906), so the two distributions are not separable.
3. **`cross-lingual` is the lowest of every type** (best relevant p10 0.8808) against
   `semantic` at 0.9208 — a threshold tuned on the aggregate would cut cross-lingual first.

## The threshold curve is the answer

| threshold | raw cosine | answerable hit | answerable full recall | unanswerable abstain |
| --- | --- | --- | --- | --- |
| 0.875 | 0.75 | 1.0000 | 1.0000 | 0.0000 |
| 0.900 | 0.80 | 0.8947 | 0.8947 | 0.2000 |
| 0.925 | 0.85 | 0.6316 | 0.6316 | 0.4000 |
| 0.950 | 0.900 | 0.1579 | 0.1579 | 1.0000 |

**Abstention never rises without full recall falling.** Every threshold that refuses an
unanswerable question refuses required relevant passages at the same rate. There is no
operating point that buys the first without paying the second.

So the conclusion is not "sweep 0.7–0.8 instead of 0–0.6". It is: **a single dense
similarity threshold cannot carry both recall and abstention on this corpus**, and the next
mechanism to evaluate is a different signal — reranker score, top1−top2 margin, per-query
thresholds, or claim-level answerability. That is a much more useful finding than a best
threshold would have been, and it is why this PR was worth doing before more clusters.

## What this PR does not do

- It does not choose a threshold. That is now blocked on a signal that can separate the
  distributions.
- The `eval:threshold` sweep grid is left as it is. Re-gridding it would move a knob whose
  curve is flat where it matters and precipitous where it does not — the diagnostic is the
  replacement for re-gridding, not a preamble to it.

## Testing

- `npm run typecheck` — clean
- `npm test` — 504 pass
- `npm run eval:scores` — writes `docs/eval/scores-v1.6.{json,md}`
- The committed baseline is **unchanged**: `--eval-scores` is off by default, so the
  CI-diffed JSON does not grow a score series it has no use for

Part of #192 (child 2: corpus diagnostics).
…ne transform

`score = (1 + cosine) / 2` is an affine map, so an **absolute** score converts as
`cosine = 2·score − 1`. A **difference** of scores does not: the `+1` cancels and
`Δcosine = 2·Δscore`.

The margin row applied the absolute transform, which reported `2Δscore − 1` — for a
margin near zero that is a cosine margin near **−1**, a sign flip on top of a scale error:

```
                    before            after
margin p10   -0.0272 (-1.054)   -0.0272 (-0.054)
margin p50   +0.0026 (-0.995)   +0.0026 (+0.005)
```

`quantileRow` now takes the cosine transform as an argument and the margin row passes
`toCosineMargin`, so the two cannot be confused at the call site again. The report also
says which mapping applies to which row.

This does not change the separability conclusion — the p90 for the best non-relevant
passage is 0.9386 against a worst-required-relevant p10 of 0.8906, and that comparison
uses absolute scores, which were always converted correctly.

The margin itself is also now labelled in the report as an **oracle** quantity: at runtime
nothing knows which result is relevant, so it measures how much the score separates the
two, and it is not a signal the product could use.

Part of #192 (correction to the score diagnostics).
…stions

`npm run eval:paired` reports the dense ↔ hybrid pair per question and the classification
per query type. It answers the one thing the aggregate comparison could not.

The open question was: hybrid is ahead on the full set and level on `validation`, and an
average cannot say whether that is a broad small gain or a handful of rescued cases. Those
two readings imply different next steps. Child 2 of #192.

## What it found

```
validation  19 questions   improved 2  / tied 16 / regressed 1   mean ΔnDCG@10 +0.0047
test        34 questions   improved 5  / tied 29 / regressed 0   mean ΔnDCG@10 +0.0432
```

**The gain is not broad — it is five questions on `test` and two on `validation`.** On the
reporting side, 29 of 34 questions are untouched. The mean ΔnDCG@10 of +0.0432 is carried by
`q013` (rank 11 → 4, ΔnDCG +0.43), `q059` (5 → 2), `q026` (2 → 1), `q049` (3 → 2) and `q052`
(2 → 1).

## And it is not the story we had started telling ourselves

The plausible story was "hybrid helps the hard negatives". The breakdown says otherwise:

| type | n (test) | improved | tied | regressed |
| --- | --- | --- | --- | --- |
| cross-lingual | 5 | **0** | 5 | 0 |
| hard-negative | 5 | **1** | 4 | 0 |
| multi-hop | 3 | 1 | 2 | 0 |
| semantic | 15 | **3** | 12 | 0 |
| exact | 4 | 0 | 4 | 0 |

Three of the five wins are `semantic`, one is `hard-negative` and one `multi-hop`. With n=1
of 5, "hybrid helps hard-negative questions" is **not** supported at this size — which is
exactly the pattern that would have been asserted from the aggregate number alone.

**`cross-lingual` is untouched on both splits: 0 improved, 0 regressed.** Adding a sparse
channel does nothing for the category that is measurably weakest, which is itself a useful
negative result and a reason not to reach for hybrid as the answer to it.

No question was found by one strategy and missed by the other, on either split, so no
regression hides behind the `0 = not found` sentinel.

## How

Metrics come from `src/main/eval/metrics.ts` — imported, not reimplemented — so a delta
cannot disagree with the metric it is a delta of. `firstRelevantRank` is `0` for "never
found", so the miss cases are classified explicitly rather than folded into an arithmetic
delta that cannot express them.

`validation` is the selecting side; the `test` breakdown **explains** the observed
difference and is labelled as not-for-selection in the report itself. Nothing here changes
the shipped strategy.

## Testing

- `npm run typecheck` — clean
- `npm run eval:paired` — 4 harness runs, writes `docs/eval/paired-v1.6.{json,md}`
- The committed baseline is unchanged

Part of #192 (child 2: corpus diagnostics).
…al PR

Closes RAG Eval v2 Phase 1. This is the final expansion of the eval system — the
conclusions below are the point of the phase, not more infrastructure.

## What was added

Fifteen unanswerable questions whose **subject is discussed at length in the corpus and
whose answer is not there** — "How much GPU memory does Gateway API v3 need?", "Which
laboratory is accredited to analyse the estuary transect samples?", "How many staff work on
the reservoir monitoring programme?".

Every specific term was checked absent before the question was written (`gpu`, `vram`, `sla`,
`uptime`, `expiry`, `sdk`, `self-hosted`, `accredit`, `vendor`, `supplier`, `funding`,
`grant`, `staff`, `warranty`, `retention`, `encrypt`, `pricing`, `invoice`, `audit` all
return nothing). That check is what makes them unanswerable rather than believed to be, and
it is the part that cannot be automated away.

`validation` unanswerable goes from **5 to 15**, so abstention resolution goes from 20 % per
question to 6.7 %. The earlier "Who won the 2018 FIFA World Cup?" questions were a sanity
check for a query that is obviously out of scope; these are the realistic shape — the user's
actual hallucination risk is asking a question about a document that *is* in the library.

## The conclusion survives the larger sample

```
unanswerable max candidate    min 0.8978 (cosine 0.796)   p50 0.9279 (0.856)   max 0.9435 (0.887)
worst required relevant       p10 0.8906 (cosine 0.781)
```

Even the **lowest** unanswerable top score is above the p10 of the worst required relevant
passage. With n=15 instead of n=5 the distributions still do not separate, and the threshold
curve is unchanged in shape:

| threshold | raw cosine | answerable full recall | unanswerable abstain |
| --- | --- | --- | --- |
| 0.875 | 0.750 | 1.0000 | 0.0000 |
| 0.900 | 0.800 | 0.8947 | 0.0667 |
| 0.925 | 0.850 | 0.6316 | 0.4000 |
| 0.950 | 0.900 | 0.1579 | 1.0000 |

Abstention still never rises without full recall falling. The finer resolution makes the
finding *more* solid, not different.

Regenerated to stay consistent: baseline, scores, threshold and sweep. `paired-v1.6` is
unchanged, because adding unanswerable questions does not touch any answerable one.

## Where Phase 1 leaves the product

1. The original 19-chunk benchmark could not guide a product decision; the current one
   (53 chunks, 78 questions, per-type metrics, held-out split) can.
2. **Cross-lingual retrieval is the largest measured gap**: Recall@5 0.6667 and nDCG@10
   0.4173 against 0.93–1.00 for same-language questions. Chinese over a Chinese source
   scores 1.0000, so it is the cross-language matching, not Chinese.
3. **Hybrid has ranking value and no held-out mandate.** Its full-set advantage is five
   questions out of 34 (`paired-v1.6`), mostly `semantic`, and it does nothing for
   cross-lingual. Dense stays the default.
4. `contextK = 3` is justified against 5 and 8 on this corpus; `candidateK` above 10 is not
   distinguishable.
5. `threshold = 0.5` means **cosine ≥ 0**, not cosine ≥ 0.5.
6. **A single embedding similarity threshold cannot decide answerability** — the relevant
   and non-relevant and unanswerable score distributions overlap, and every threshold that
   abstains also drops required evidence.

(6) is the phase's real result: it says stop tuning this knob and reach for a different
mechanism.

## Stop line

Phase 1 ends here. Not done, and deliberately not started: public benchmarks, reranker eval,
generator faithfulness, citation entailment, growing the corpus to 300 chunks, an
`Abstention Signal Evaluation`. Those build a more complete benchmark; they do not unblock
KnowNote development, and (6) already says the next useful work is a different retrieval
mechanism rather than a better measurement of this one.
@github-actions github-actions Bot added the skip-changelog Exclude from generated release notes label Oct 1, 2026
@mrsibe
mrsibe merged commit 6cf23ea into main Oct 1, 2026
4 checks passed
@mrsibe mrsibe mentioned this pull request Oct 1, 2026
6 tasks done
@mrsibe
mrsibe deleted the chore/eval-v2-closeout branch October 2, 2026 10:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

skip-changelog Exclude from generated release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Epic] Retrieval Eval v2 — Phase 1 closeout

1 participant