Skip to content

feat(eval): measure the context window, and name what still needs a model - #198

Merged
mrsibe merged 1 commit into
feat/eval-saturation-rulefrom
feat/eval-context-metrics
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-saturation-rulefrom
feat/eval-context-metrics

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #197 (feat/eval-saturation-rule). Retarget to main after that merges.

What does this PR do?

Measures the window the prompt actually receives, and states plainly what still needs a model. Child 5 of #192 — the deterministic half.

The gap

The harness measured the retriever but never the context window. evidenceK: 5 was a separate constant from the contextK: 3 production uses, so the one metric that looked at a window looked at a different one than the product does.

The change

contextK is now the window for both context metrics:

  • contextPrecision@contextK — of the first contextK passages, the share covering ground truth. This is what evidencePrecisionAt5 was, at the production width.
  • contextRecall@contextK — the share of needed ground-truth blocks that made it into that window. Distinct from Recall@10: a block found at rank 4 is invisible when contextK = 3, and that is a product fact, not a ranking fact.

Both are deterministic: the dataset says which blocks answer the question, so no model is needed to score a window. Together they are the trade-off a contextK decision makes — a wider window finds more and carries more noise — which is what the sweep in the next child needs.

Baseline: contextPrecision@3 = 0.3556, contextRecall@3 = 1.0000. The needed evidence is always inside the top 3 on this corpus, but only about a third of what is inside the window is relevant. The second number is the one that says the window is paying for passages that do not answer the question.

evidenceK is removed from the harness config, so there is exactly one context width.

What is deliberately not here

Faithfulness, completeness, answer correctness and noise sensitivity are not added. They need a generative model, and the harness runs offline with only the pinned embedding model — the same constraint that keeps the reranker unmeasured (#170). This PR does not fake them. The report's Definitions section now says so explicitly, and the experiment script's "Not evaluated" note points at the same constraint.

Adding an LLM judge is a separate change that has to solve model pinning first. It is not a line of code, and shipping a judge that silently degrades to "no metrics" in CI would be worse than shipping none.

Testing

  • npm run typecheck — clean
  • npm test — 497 pass
  • npm run eval twice — byte-identical docs/eval/baseline-v1.6.json
  • npm run eval:retrieval — regenerated; hybrid still clears the amended rule

Related

Part of #192. Child 5 (deterministic half: context precision/recall).

…odel

The harness measured the retriever but never the window the prompt actually gets.
`evidenceK: 5` was a separate constant from the `contextK: 3` production uses, so the
one metric that looked at a window looked at a different one than the product does.
Child 5 of #192.

**`contextK` is now the window for both context metrics.**

- `contextPrecision@contextK` — of the first `contextK` passages, the share covering
  ground truth. This is what `evidencePrecisionAt5` was, at the production width.
- `contextRecall@contextK` — the share of needed ground-truth blocks that made it into
  that window. Distinct from `Recall@10`: a block found at rank 4 is invisible when
  `contextK = 3`, and that is a product fact, not a ranking fact.

Both are **deterministic**: the dataset says which blocks answer the question, so no
model is needed to score a window. Together they are the trade-off a `contextK`
decision makes — a wider window finds more and carries more noise — which is what the
sweep in the next child needs.

Baseline: `contextPrecision@3 = 0.3556`, `contextRecall@3 = 1.0000` — the needed
evidence is always inside the top 3 on this corpus, but only about a third of what is
inside the window is relevant. That second number is the one that says the window is
paying for passages that do not answer the question.

**What is deliberately not here.** Faithfulness, completeness, answer correctness and
noise sensitivity need a generative model. The harness runs offline with only the
pinned embedding model — the same constraint that keeps the reranker unmeasured
(#170) — so this PR does not add them and does not fake them: the report's
Definitions section now says so, and the "Not evaluated" line in
`eval-retrieval.mjs` points at the same constraint. Adding an LLM judge is a separate
change that has to solve model pinning first, not a line of code.

`evidenceK` is removed from the harness config, so there is exactly one context width.

## Testing

- `npm run typecheck` — clean
- `npm test` — 497 pass
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`
- `npm run eval:retrieval` — regenerated; hybrid still clears the amended rule

Part of #192 (child 5, deterministic half).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit 7713b40 into feat/eval-saturation-rule Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant