Skip to content

feat(eval): confusion clusters, and the corpus reaches 53 chunks - #205

Merged
mrsibe merged 2 commits into
fix/eval-review-followupsfrom
feat/eval-corpus-hard-negatives
Sep 30, 2026
Merged

mrsibe merged 2 commits into
fix/eval-review-followupsfrom
feat/eval-corpus-hard-negatives

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #204 (fix/eval-review-followups). Retarget to main after the chain merges.

What does this PR do?

The first two confusion clusters of corpus v2 — deliberately confusable material, not "near-duplicate documents that lower Recall@5". The goal is a benchmark that can tell retrieval strategies apart at all, and lowering Recall is a consequence of that rather than the objective.

The clusters

1. Gateway API v1 / v2 / v3 + a migration guide. Four documents sharing almost all their vocabulary and differing in every number:

v1 v2 v3
Default timeout 30000 ms 60000 ms 45000 ms
Recommended retries 2 5 3
Backoff base 500 ms 2000 ms 1000 ms
Rate limit 600 rpm 3000 rpm 1200 rpm
Context window 8192 tokens 32768 tokens 65536 tokens

All four discuss timeout, retry, backoff, rate limit, context window in that order. The migration guide restates every number in comparison tables, so it contains v1's timeout and v2's and v3's in a single block — the hardest negative in the set.

2. Monitoring protocols. Four documents that are confusable with each other and with the river-monitoring.md / lake-monitoring.md already in the corpus:

reservoir coastal estuary groundwater
Sampling interval fortnightly hourly daily monthly
Replicates 4 3 5 2
Sensor depth 5 m below surface 1 m below surface 2 m below surface 15 m below water table

Each states the network's extreme value ("the least frequent cadence in the network", "the highest of any programme"), so several passages read as plausible and only one is right.

Corpus: 19 → 53 chunks, 13 → 21 documents, 44 → 63 questions.

The tool that made it possible

npm run eval:blocks <file.md> prints the real document_blocks.order values through the same loader and block builder the harness uses. Ground truth block is a pipeline ordinal, not a line number a human counted; authoring a corpus by guessing it is how a dataset quietly drifts. eval/README.md now says so and points at the command.

What it measured

                 before    after
chunks           19        53
Recall@5         0.9211    0.8774
nDCG@10          0.8476    0.7804
MAP@10           0.7965    0.7280
type n Recall@5 nDCG@10
cross-lingual 9 0.6667 0.4173
multi-hop 5 0.7000 0.7218
semantic 21 0.9048 0.8542
exact 8 1.0000 0.8663
hard-negative 7 1.0000 0.8758
zh 3 1.0000 1.0000

hard-negative has perfect Recall@5 and a much worse nDCG — the signature the material was built for: the correct passage is in the top 5, and a confusable sibling keeps outranking it. That is a ranking problem the strategy comparison can now see, and it is invisible on a corpus where the right answer is always at rank 1.

multi-hop fell to 0.7000: the comparison questions need four locations each and do not get all four.

A result worth pausing on

The sweep (all 63 questions) now shows hybrid ahead of dense on the deciding metric:

Recall@5 nDCG@10
dense, candidateK=20 0.8774 0.7804
hybrid, candidateK=20 0.8962 0.8098
hybrid, candidateK=5 0.9104 0.8101

But on the validation split the same comparison is a dead heat (both 0.8684), so the held-out rule still declines to adopt and test reports dense only.

That discrepancy is the honest state of things, and it is a property of the split, not of hybrid: 24 validation questions is not enough to see a gain the full set shows. The conclusion is to grow the validation split, not to go back to selecting on everything — which is exactly what produced the earlier, over-confident "hybrid clears the rule".

Unanswerable is still 0/10

meanCandidatesRetrieved sits at the candidateK cap and meanContextPassages at contextK for every unanswerable question, so abstention remains unmeasurable here. Two of the ten are near-miss questions — "How much GPU memory does Gateway API v2 require?" and "How many litres per second does the reservoir release downstream?" — about subjects the corpus discusses at length and never answers. Those are the realistic hallucination shape; the FIFA questions were only ever a sanity check.

Scope, honestly

53 chunks, not 150–300. Two clusters is not the target either. candidateK ∈ {5,10,20,40} has more room than it did at 19 chunks, but candidateK ≥ 10 still behaves identically, so the axis is only partly unlocked and threshold remains unmeasurable.

The remaining clusters are the same exercise and the same verification loop. This PR stops where the verified work stops.

Testing

  • npm test — 504 pass
  • npm run eval twice — byte-identical baseline (63 questions, 53 chunks)
  • npm run eval:retrieval, eval:threshold, eval:sweep — all regenerated
  • eval/splits.json extended: 24 validation / 39 test, stratified by type

Related

Part of #192. Child 2, stage 3 (confusion clusters).

@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
…egatives

The first of the confusion clusters corpus v2 is built from. Not "near-duplicate documents
that lower Recall@5" — deliberately confusable material, so the benchmark can tell retrieval
strategies apart at all.

## The cluster

Four documents that share almost all their vocabulary and differ in every number:

| | v1 | v2 | v3 |
| --- | --- | --- | --- |
| Path | `/v1/complete` | `/v2/generate` | `/v3/chat` |
| Default timeout | 30000 ms | 60000 ms | 45000 ms |
| Recommended retries | 2 | 5 | 3 |
| Backoff base | 500 ms | 2000 ms | 1000 ms |
| Rate limit | 600 rpm | 3000 rpm | 1200 rpm |
| Context window | 8192 tokens | 32768 tokens | 65536 tokens |

All four documents discuss *timeout, retry, backoff, rate limit and context window* in that
order, so a query naming one field matches four passages and only one is right. A migration
guide restates every number in comparison tables, which is the hardest negative of the set:
it contains v1's timeout **and** v2's **and** v3's in one block.

## The tool that made it possible

`npm run eval:blocks <file.md>` prints the real `document_blocks.order` values through the
same loader and block builder the harness uses. Ground truth `block` is a pipeline ordinal,
not a line number a human counted; authoring a corpus by guessing it is how a dataset
quietly drifts. `eval/README.md` now says so and points at the command.

## What it measured

```
                 before    after
chunks           19        35
Recall@5         0.9211    0.8913
nDCG@10          0.8476    0.7915
MAP@10           0.7965    0.7361
```

By type, the new cluster behaves as designed:

| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| hard-negative | 4 | **1.0000** | **0.7827** |
| multi-hop | 4 | 0.7500 | 0.7289 |
| semantic | 19 | 0.9474 | 0.8974 |
| cross-lingual | 9 | 0.6667 | 0.4311 |

`hard-negative` has perfect Recall@5 and a poor nDCG, which is the exact signature of this
kind of material: the correct passage **is** in the top 5, but a confusable sibling outranks
it. That is a ranking problem the strategy comparison can now see, and it is invisible on a
corpus where the right answer is always at rank 1.

`multi-hop` fell to 0.75 because the two comparison questions need four locations each and
do not get all four.

Also added: 8 questions in total (2 exact, 4 hard-negative, 1 semantic, 2 multi-hop, and two
**near-miss unanswerable** — "How much GPU memory does Gateway API v2 require?" and "What is
the monthly subscription price of Gateway API v3?", both about documents that discuss the
subject at length and never mention the answer). The FIFA questions were a sanity check for
a totally out-of-scope query; these are the ones that actually test abstention, because the
subject *is* in the library.

Unanswerable abstention is still 0/8, with `meanCandidatesRetrieved` at the `candidateK`
cap of 20 and `meanContextPassages` at 3.

## Scope, honestly

**35 chunks, not 150–300.** One cluster is not the target. The mechanism is proven and
verified end to end, and the remaining clusters are the same exercise repeated — but this PR
does not pretend the index is large enough for `candidateK ∈ {5,10,20,40}` to be a real
sweep. That arrives with the remaining clusters.

Part of #192 (child 2, stage 3: first confusion cluster).
Four more deliberately confusable documents, this time in the monitoring family the corpus
already had two members of, so the cluster is confusable with the *existing* documents as
well as internally.

| | reservoir | coastal | estuary | groundwater |
| --- | --- | --- | --- | --- |
| Sampling interval | fortnightly | hourly | daily | monthly |
| Replicates | 4 | 3 | 5 | 2 |
| Sensor depth | 5 m below surface | 1 m below surface | 2 m below surface | 15 m below water table |

Every document discusses *sampling interval, replicate samples and sensor depth* in the same
order, and each states the network's lowest or highest value ("the least frequent cadence in
the network", "the highest of any programme"), so a query lands on several plausible passages
and only one is right.

Corpus: 19 → 53 chunks, 13 → 21 documents, 54 → 63 questions.

## What it measured

```
                 after cluster 1   after cluster 2
chunks           35                53
Recall@5         0.8913            0.8774
nDCG@10          0.7915            0.7804
MAP@10           0.7361            0.7280
```

| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| cross-lingual | 9 | 0.6667 | 0.4173 |
| multi-hop | 5 | 0.7000 | 0.7218 |
| semantic | 21 | 0.9048 | 0.8542 |
| exact | 8 | 1.0000 | 0.8663 |
| hard-negative | 7 | **1.0000** | **0.8758** |
| zh | 3 | 1.0000 | 1.0000 |

`hard-negative` keeps the signature the cluster was built for: the correct passage is always
in the top 5 and a confusable sibling keeps outranking it.

## A result worth pausing on

The sweep (all 63 questions) now shows **hybrid ahead of dense on the deciding metric**:

| | Recall@5 | nDCG@10 |
| --- | --- | --- |
| dense, candidateK=20 | 0.8774 | 0.7804 |
| hybrid, candidateK=20 | **0.8962** | **0.8098** |
| hybrid, candidateK=5 | **0.9104** | 0.8101 |

But on the **validation** split the same comparison is a dead heat (both 0.8684), so the
held-out rule still declines to adopt.

That discrepancy is the honest state of things, and it is a property of the split rather
than of hybrid: 24 validation questions is not enough to see a gain that the full set shows.
The conclusion is to grow the validation split, **not** to go back to selecting on
everything — which is what produced the earlier, over-confident "hybrid clears the rule".

## Unanswerable is still 0/10

`meanCandidatesRetrieved` sits at the `candidateK` cap and `meanContextPassages` at
`contextK` for every unanswerable question, so abstention remains unmeasurable here. Two of
the ten are near-miss questions ("How much GPU memory does Gateway API v2 require?", "How
many litres per second does the reservoir release downstream?") about subjects the corpus
discusses at length — the realistic hallucination shape rather than the FIFA sanity check.

## Scope, honestly

**53 chunks, not 150–300.** Two clusters is not the target either. `candidateK ∈
{5,10,20,40}` now has more room than it did at 19 chunks, but `candidateK ≥ 10` still
behaves identically, so the axis is only partly unlocked. The remaining clusters are the same
exercise and the same verification loop; this PR stops where the verified work stops.

Part of #192 (child 2, stage 3: confusion clusters).
@mrsibe
mrsibe force-pushed the fix/eval-review-followups branch from 22f070c to 84e9e0a Compare September 30, 2026 10:07
@mrsibe
mrsibe force-pushed the feat/eval-corpus-hard-negatives branch from d171f8f to 3ffee33 Compare September 30, 2026 10:07
@mrsibe
mrsibe merged commit af481bd into fix/eval-review-followups Sep 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant