Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
25e1ada
fix(eval): run the harness on Windows and keep the baseline POSIX
mrsibe Sep 30, 2026
a444a86
feat(eval): separate candidateK from contextK, and run the production…
mrsibe Sep 30, 2026
54c892f
feat(eval): hit rate, MAP, and per-query-type metrics
mrsibe Sep 30, 2026
a2ccf05
feat(eval): derive the similarity threshold on validation, report on …
mrsibe Sep 30, 2026
a52b9b7
feat(eval): adopt on the first metric with headroom, not a saturated one
mrsibe Sep 30, 2026
b77257a
feat(eval): measure the context window, and name what still needs a m…
mrsibe Sep 30, 2026
27fdda2
feat(eval): bounded parameter sweep with a trade-off dashboard
mrsibe Sep 30, 2026
aee2f89
fix(eval): refuse a context window wider than the retrieval depth
mrsibe Sep 30, 2026
d452463
feat(eval): unanswerable questions, and an explicit split manifest
mrsibe Sep 30, 2026
3f60d7c
feat(eval): nine cross-lingual questions, and the corpus stops being …
mrsibe Sep 30, 2026
b1bdb0b
fix(retrieval): trace the dense channel's threshold for hybrid too
mrsibe Sep 30, 2026
84e9e0a
fix(eval): select the strategy on validation, and stop calling candid…
mrsibe Sep 30, 2026
20a8c75
feat(eval): a confusion cluster, and the corpus gets its first hard n…
mrsibe Sep 30, 2026
3ffee33
feat(eval): a second confusion cluster, and the corpus reaches 53 chunks
mrsibe Sep 30, 2026
08fa7c1
feat(eval): dense score diagnostics, and the answer is not a threshold
mrsibe Sep 30, 2026
036a95a
fix(eval): convert a score margin with ×2, not with the absolute affi…
mrsibe Sep 30, 2026
8a99107
feat(eval): per-question paired deltas, and hybrid's gain is five que…
mrsibe Sep 30, 2026
d368832
feat(eval): fifteen near-miss unanswerable questions, and the last ev…
mrsibe Sep 30, 2026
a850f9f
chore(eval): integrate retrieval v2 and close Phase 1 scope
mrsibe Oct 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions .prettierignore
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,24 @@ docs/eval/baseline-*.md
docs/eval/chunking-*.json
docs/eval/chunking-*.md

# Same reason again: generated by `scripts/eval-threshold.mjs`. Regenerate with
# `npm run eval:threshold`.
docs/eval/threshold-*.json
docs/eval/threshold-*.md

# Same reason again: generated by `scripts/eval-sweep.mjs`. Regenerate with
# `npm run eval:sweep`.
docs/eval/sweep-*.json
docs/eval/sweep-*.md

# And by `scripts/eval-scores.mjs`. Regenerate with `npm run eval:scores`.
docs/eval/scores-*.json
docs/eval/scores-*.md

# And by `scripts/eval-paired.mjs`. Regenerate with `npm run eval:paired`.
docs/eval/paired-*.json
docs/eval/paired-*.md

# And the same again for `scripts/eval-retrieval.mjs` (#77).
docs/eval/retrieval-*.json
docs/eval/retrieval-*.md
61 changes: 61 additions & 0 deletions docs/eval/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# Eval closeout — v1.6

This directory holds the retrieval eval reports. The versioned report `*.json` / `*.md` files here are
harness output — regenerate it with the script named in its header, never edit it by
hand. `eval/README.md` documents the harness itself.

## What this closeout lands

The eval work reaches v1.6 and stops at the retrieval layer:

- **Two Ks.** `candidateK` is the first-stage width per channel, `contextK` is how many
passages reach the prompt. `baseline-v1.6` records both.
- **An explicit split.** `eval/splits.json` assigns every question to `validation` or
`test`. A parameter is selected on `validation` and reported on `test`.
- **Shipped metrics.** Hit rate@5, MAP@10, context precision/recall, per-query-type
breakdown, and a separate unanswerable group.
- **Experiments, not wins.** `retrieval-v1.6`, `threshold-v1.6`, `sweep-v1.6`,
`scores-v1.6` and `paired-v1.6` compare strategies and parameters against the
baseline.

## What did **not** change

Production retrieval is unchanged: **dense**, `candidateK = 20`, `contextK = 3`,
`threshold = 0.5` (`src/main/ipc/chatHandlers.ts`). No strategy cleared the adoption
rule on `validation`, so `retrieval-v1.6` reports the shipped dense strategy and no
winner. The threshold sweep is **flat** — every threshold from 0 to 0.6 behaves
identically on this corpus — so there is no evidence to move `threshold` either. These
are negative results: validation does not authorize changing the defaults. They do
not establish that other strategies have no value.

## Corpus state

21 documents, 53 chunks, 78 questions (53 answerable, 25 unanswerable). The unanswerable
group is what a threshold decision should move; on this corpus it does not, because no
candidate is ever filtered out.

## Phase 1 boundary

Retrieval Eval v2 Phase 1 (#192) closes with this integration. No further K,
threshold, hybrid or retrieval-metric tuning is planned under that epic.

Generator evaluation is tracked in #213 and citation evaluation in #214. Public
QASPER/MIRACL-zh subsets, a pinned offline reranker (historical PR #170), further
corpus scaling and token accounting beyond the character proxy are deferred, not
claimed complete and not prerequisites for this phase's closure.

The original ~300-chunk target is unmet. At 53 chunks, only 19 answerable questions
are in `validation`: this is a small diagnostic benchmark, not evidence of broad
real-world performance. Expanding it is explicitly outside this closeout.

## Regenerate

```bash
npm run eval:prepare # one-time, networked: pin the embedding model
npm run eval # baseline-v1.6.{json,md}
npm run eval:retrieval # retrieval-v1.6.{json,md}
npm run eval:threshold # threshold-v1.6.{json,md}
npm run eval:sweep # sweep-v1.6.{json,md}
npm run eval:scores # scores-v1.6.{json,md}
npm run eval:paired # paired-v1.6.{json,md}
```
Loading
Loading