Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions docs/eval/baseline-v1.6.json
Original file line number Diff line number Diff line change
Expand Up @@ -110,9 +110,10 @@
],
"unanswerable": {
"questions": 6,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19
"abstentionCount": 0,
"retrievalAbstentionRate": 0,
"meanCandidatesRetrieved": 19,
"meanContextPassages": 3
},
"perQuestion": [
{
Expand Down
26 changes: 18 additions & 8 deletions docs/eval/baseline-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,20 +45,30 @@ The type comes from `type` in `questions.jsonl`; untagged questions report as

### Unanswerable questions

These carry no ground truth, so the correct outcome is that retrieval finds nothing. They
These carry no ground truth, so the correct outcome is that retrieval returns nothing. They
are excluded from every metric above — a missing ground truth is not a miss — and reported
here instead. A higher **no-results** rate is better on this row, which is the opposite of
how it reads everywhere else, and `mean passages retrieved` is how much irrelevant context
was pulled in anyway. This is the row a threshold decision should move.
here instead.

**This measures retrieval-level abstention, not the model refusing.** No generator runs in
this harness, so it can show that no candidate passed the threshold; it cannot show that the
final answer would say "not in your sources". A true system refusal rate needs a
generator eval.

A higher abstention rate is better on this row, the opposite of how every other row reads.
The two sizes are kept apart on purpose: **candidates passing the threshold** can be as high
as `candidateK` (the harness fetches that many to compute `Recall@10`), while **passages in
the context window** is what a user's prompt would actually receive. A large first number
with a small second one means the threshold filters nothing and the window is all noise.

| Metric | Value |
| --- | --- |
| Unanswerable questions | 6 |
| Returned no results | 0.0000 (0/6) |
| Mean passages retrieved | 19.00 |
| Retrieval abstained | 0.0000 (0/6) |
| Mean candidates passing the threshold | 19.00 |
| Mean passages in the context window | 3.00 |

Timing is informational only and is **not** frozen: indexing 1552 ms, query
p50 11.82 ms, p95 14.45 ms on the
Timing is informational only and is **not** frozen: indexing 1583 ms, query
p50 13.38 ms, p95 16.66 ms on the
machine that produced this file. Timing and index size depend on hardware and on the
corpus, so they must never be the reason two runs differ.

Expand Down
113 changes: 70 additions & 43 deletions docs/eval/retrieval-v1.6.json
Original file line number Diff line number Diff line change
@@ -1,66 +1,93 @@
{
"baseline": "dense",
"chunking": "1000/100",
"strategies": [
"baseline": "v1.6",
"split": {
"selects": "validation",
"reports": "test",
"manifest": "eval/splits.json"
},
"decision": {
"primary": "recallAt5",
"saturated": [],
"winner": null
},
"validation": [
{
"id": "dense",
"label": "dense (vector)",
"split": "validation",
"questions": 16,
"answerableCount": 13,
"chunking": "1000/100",
"chunkCount": 19,
"split": "all",
"questions": 44,
"answerableCount": 38,
"recallAt1": 0.657895,
"recallAt5": 0.921053,
"recallAt1": 0.576923,
"recallAt5": 0.923077,
"recallAt10": 1,
"mrr": 0.800909,
"ndcgAt10": 0.84764,
"hitRateAt5": 0.921053,
"mapAt10": 0.796523,
"contextPrecision": 0.324561,
"contextRecall": 0.921053,
"indexingMs": 1532,
"latencyP95Ms": 14.38
"mrr": 0.778846,
"ndcgAt10": 0.827608,
"hitRateAt5": 0.923077,
"mapAt10": 0.766026,
"contextPrecision": 0.333333,
"contextRecall": 0.923077,
"latencyP95Ms": 16.02
},
{
"id": "sparse",
"label": "sparse (BM25)",
"split": "validation",
"questions": 16,
"answerableCount": 13,
"chunking": "1000/100",
"chunkCount": 19,
"split": "all",
"questions": 44,
"answerableCount": 38,
"recallAt1": 0.578947,
"recallAt5": 0.684211,
"recallAt10": 0.684211,
"mrr": 0.635965,
"ndcgAt10": 0.64607,
"hitRateAt5": 0.684211,
"mapAt10": 0.631579,
"contextPrecision": 0.245614,
"contextRecall": 0.684211,
"indexingMs": 1529,
"latencyP95Ms": 1.85
"recallAt1": 0.576923,
"recallAt5": 0.615385,
"recallAt10": 0.615385,
"mrr": 0.615385,
"ndcgAt10": 0.609209,
"hitRateAt5": 0.615385,
"mapAt10": 0.602564,
"contextPrecision": 0.230769,
"contextRecall": 0.615385,
"latencyP95Ms": 2.96
},
{
"id": "hybrid",
"label": "hybrid (RRF of dense + BM25)",
"split": "validation",
"questions": 16,
"answerableCount": 13,
"chunking": "1000/100",
"chunkCount": 19,
"split": "all",
"questions": 44,
"answerableCount": 38,
"recallAt1": 0.684211,
"recallAt5": 0.921053,
"recallAt1": 0.576923,
"recallAt5": 0.923077,
"recallAt10": 1,
"mrr": 0.814066,
"ndcgAt10": 0.857352,
"hitRateAt5": 0.921053,
"mapAt10": 0.80968,
"contextPrecision": 0.324561,
"contextRecall": 0.921053,
"indexingMs": 1579,
"latencyP95Ms": 17.04
"mrr": 0.778846,
"ndcgAt10": 0.827608,
"hitRateAt5": 0.923077,
"mapAt10": 0.766026,
"contextPrecision": 0.333333,
"contextRecall": 0.923077,
"latencyP95Ms": 19.47
}
],
"test": [
{
"id": "dense",
"label": "dense (vector)",
"split": "test",
"questions": 28,
"answerableCount": 25,
"chunking": "1000/100",
"chunkCount": 19,
"recallAt1": 0.7,
"recallAt5": 0.92,
"recallAt10": 1,
"mrr": 0.812381,
"ndcgAt10": 0.858056,
"hitRateAt5": 0.92,
"mapAt10": 0.812381,
"contextPrecision": 0.32,
"contextRecall": 0.92,
"latencyP95Ms": 14.49
}
]
}
43 changes: 27 additions & 16 deletions docs/eval/retrieval-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,15 +4,27 @@ Generated by `node scripts/eval-retrieval.mjs`. Numbers are harness output; do n

## What was measured

Every strategy runs the real RAG eval harness against the same corpus and the same questions
as `baseline-v1.6.json` (split `all`, 44 questions of which
38 are answerable), with chunking held fixed at 1000/100. Only the retrieval strategy changes.
Each strategy runs the real harness on the **validation** split of
`eval/splits.json`, with chunking held fixed at 1000/100.

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | MAP@10 | Context P | Query p95 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| dense (vector) | 0.6579 | 0.9211 | 0.8009 | 0.8476 | 0.7965 | 0.3246 | 14.38 ms |
| sparse (BM25) | 0.5789 | 0.6842 | 0.6360 | 0.6461 | 0.6316 | 0.2456 | 1.85 ms |
| hybrid (RRF of dense + BM25) | 0.6842 | 0.9211 | 0.8141 | 0.8574 | 0.8097 | 0.3246 | 17.04 ms |
The strategy is then chosen **there**, and only the chosen one (plus the shipped default)
is re-run on **test**. The choice never sees `test`; `test` only reports. The previous
version decided on `split = all`, which scored the choice on the questions it was fitted
to.

## Validation — this is where the choice happens

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | MAP@10 | Context P | n | p95 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| dense (vector) | 0.5769 | 0.9231 | 0.7788 | 0.8276 | 0.7660 | 0.3333 | 16 | 16.02 ms |
| sparse (BM25) | 0.5769 | 0.6154 | 0.6154 | 0.6092 | 0.6026 | 0.2308 | 16 | 2.96 ms |
| hybrid (RRF of dense + BM25) | 0.5769 | 0.9231 | 0.7788 | 0.8276 | 0.7660 | 0.3333 | 16 | 19.47 ms |

## Test — reported, not selected

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | MAP@10 | Context P | n | p95 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| dense (vector) | 0.7000 | 0.9200 | 0.8124 | 0.8581 | 0.8124 | 0.3200 | 28 | 14.49 ms |

## Not evaluated

Expand All @@ -21,24 +33,23 @@ model is not available offline and inventing its numbers would defeat the point
the harness. It stays open until a model can be pinned the way the embedding model
is.

## Saturation

No metric in the rule is saturated on this corpus. The deciding metric on this corpus is `recallAt5`. A saturated metric is still reported, because "this corpus cannot move it" is itself
information; it is just not allowed to decide the comparison.

## Adoption rule

> Adopt a strategy when it improves the **first metric with headroom** — in the order
> `recallAt5`, `nDCG@10`, `MRR`, `MAP@10` — and regresses none of the others. A metric
> already at its maximum has no headroom and cannot decide anything; a rule that depends
> on one is unsatisfiable, not strict (#192 child 10).
>
> A change that trades a large latency increase for a marginal quality gain is a product
> decision, not an automatic win.
> The rule is applied on `validation`. A change that trades a large latency increase for a
> marginal quality gain is a product decision, not an automatic win.

## Outcome

No strategy cleared the rule. The deciding metric was `recallAt5` (dense 0.9211); the strategies either failed to improve it or regressed another metric. **Dense stays the default.** A negative result is the point of the experiment: it is the measurement that says the extra machinery is not worth its cost on this corpus, not a failure to deliver. No metric in the rule is saturated on this corpus.
**No strategy cleared the rule on validation**, so there is no adoption candidate and
`test` reports the shipped strategy only. No metric in the rule is saturated on the validation split.

A negative result is the point of the experiment: it is the measurement that says the extra
machinery is not worth its cost on this corpus, not a failure to deliver.

## Reproduce

Expand Down
Loading
Loading