Skip to content

[Epic] RAG Evaluation v2 — make the eval harness represent the product #192

Description

@mrsibe

Theme

One sentence: turn the retrieval regression test into a real RAG benchmark — one that runs the production configuration, measures the generator and the citations, and can actually discriminate between candidate algorithms.

The harness we built in v1.4/v1.5 is not the problem. dense + BM25 + RRF, Recall@K / MRR / nDCG@10, p50/p95 latency, a chunking A/B test, a pinned embedding model, and a reproducible offline run are already the hard part of the work. What we do not yet have is an evaluation whose numbers can be trusted to represent the product.

This epic is about fixing the conditions of the experiment, not replacing the algorithm.

Why the current numbers do not mean what they look like

1. The benchmark is saturated

eval/corpus/ is 13 documents / 19 chunks / 30 questions, evaluated at topK = 10. Retrieving 10 of 19 chunks per question means the harness hands the ranker ~53% of the entire knowledge base before scoring. docs/eval/retrieval-v1.5.md already says it plainly:

The rule's Recall@5 condition is saturated on this corpus: dense already scores 1.0000, so no strategy can improve it and the rule can therefore never be met here.

So the v1.5 result — hybrid better on Recall@1 (0.8667 vs 0.8333), MRR (0.9444 vs 0.9278) and nDCG@10 (0.9561 vs 0.9437) — is recorded as inconclusive not because hybrid is not better, but because the adoption rule ("Recall@5 must improve") cannot move on a corpus where Recall@5 is already perfect. Any future retrieval change is blocked by the same rule. That is a benchmark defect, not a retrieval result.

2. The harness does not run the product

Retrieval topK threshold
Production (chatHandlers.ts) dense 3 0.5
Eval (eval/run.ts) dense 10 0

We are benchmarking a retriever the user never runs. A chunk at rank 4 with cosine 0.47 is a hit in the eval and does not exist in production. Recall@5 = 100% can be reported while the shipped path silently drops the evidence. This mismatch is more consequential than the value of topK itself.

3. Candidate K and context K are conflated

There is one topK in the contract. It is doing two different jobs:

  • how wide the first stage searches (recall-oriented), and
  • how much text is assembled into the prompt (precision/budget-oriented).

HybridRetriever.ts fuses and then slices at the same number:

hits = rrfFuse([denseHits, sparseHits]).slice(0, topK)

With topK = 3 in chat, dense contributes 3 and BM25 contributes 3 — the fusion pool is at most 6, then truncated back to 3. That is not a two-stage retrieval; it is a single narrow search with fusion applied inside it.

The standard shape separates these: retrieve a wide candidate pool, then narrow it.

                  ┌→ Dense  candidateK=20 ─┐
Query → Retrieval │                        ├→ RRF → rerank → contextK=3..8 → LLM
                  └→ BM25   candidateK=20 ─┘

NVIDIA's RAG blueprint sets VDB top-k at 100 and reranker top-k at 10 for exactly this reason; the Datawhale advanced-RAG path retrieves top-20 then reranks to top-5.

4. Two things are simply not evaluated

docs/eval/retrieval-v1.5.md records that reranking is not evaluated (no offline cross-encoder could be pinned), and the harness stops at the retriever. There is no generator metric and no citation metric — while citation is the product's core differentiator and has shipped since #155/#156/#162/#164.

Scope

A. Evaluation dataset

Two suites, deliberately separate.

A1. In-domain — KnowNote Eval (the one that sets production defaults)

Public datasets test their own distribution, not the one users upload. Target scale, staged:

stage 1:  ~20 documents / ~300 chunks  / ~80 questions
stage 2:  50-100 docs   / 500-3000 chunks / 200-500 questions

Source types: PDF, Markdown, Web, DOCX, PPTX. Query types deliberately covered, because these are what separate algorithms:

exact keyword · semantic · paraphrase · single-hop · multi-chunk
multi-document · entity/number/version · hard negative · unanswerable
source-scoped · citation-required

Hard negatives are mandatory — e.g. two chunks for "GPT-4 context window" and "GPT-4 Turbo context window" where dense cross-contaminates. Unanswerable questions are mandatory: today the harness has no way to measure whether we say "not in your sources".

Ground truth extends the existing relevant[] schema with graded relevance, a reference answer, and claims, so nDCG can use graded gains and the generator/citation layers have a target:

{
  "id": "kn-001",
  "question": "...",
  "answerable": true,
  "type": "multi-source",
  "referenceAnswer": "...",
  "claims": ["...", "..."],
  "evidence": [
    { "document": "...", "page": 12, "block": 21, "quote": "...", "relevance": 3 }
  ]
}

relevance: 3 essential / 2 useful / 1 supporting / 0 irrelevant.

Ground truth must stay quote-and-location based (block/page), not chunk-id based — otherwise re-chunking invalidates the dataset and chunk-size experiments become unmeasurable. The chunk↔block mapping from #108 is what makes this work.

LLM-authored questions are allowed as candidates but must be reviewed, and the hard categories (paraphrase, hard negative, multi-hop, multi-source, unanswerable) authored deliberately. A fully synthetic set reproduces the current failure: the source sentence and the generated question share vocabulary, dense scores 1.0, and we are back to a saturated benchmark.

A2. Public — sanity and comparability

Kept small and pinned; a reproducible subset is enough. Priority order:

Dataset Tests Why
QASPER long-doc QA + supporting evidence closest to "upload a PDF, ask, cite the paragraph"; 55.5% of questions need multiple paragraphs
MIRACL-zh multilingual retrieval official human-judged zh subset; exercises multilingual-e5-small vs BM25 vs hybrid
BEIR / SciFact standard IR external sanity check that our retriever is a normal retriever
HAGRID answer + attribution reference for citation evaluation
CRUD-RAG Chinese end-to-end RAG zh generator-side reference
RAGBench end-to-end RAG relevance / adherence / completeness / utilization

The public suite answers "is our retriever competitive?" The in-domain suite answers "did the product get better?" Only the second one sets defaults. Public results must always be reported separately and labelled as out-of-distribution.

B. Config parity — one config, both paths

  • Split the contract: candidateK (first stage, per channel) and contextK (assembled into the prompt). Keep threshold explicit as a separate concern.
  • The eval harness must run the production configuration as its baseline run, and any deviation must be a named experiment.
  • HybridRetriever must fuse a wide pool and truncate once, after fusion — not slice each channel to the final K.
  • Re-derive threshold from the validation set instead of hand-picking 0.5. Cosine ~0.5 has no universal meaning; it varies with model, language, query type, chunk length and domain. Sweep 0.3 / 0.4 / 0.5 / 0.6 and watch recall, precision and the no-result rate. First stage may legitimately run threshold = off and let a reranker do the filtering.

C. Retrieval metrics

Keep Recall@K, MRR, nDCG@K. Add HitRate@K, Precision@K and MAP. Report per query type, not only aggregate — a strategy that helps entity queries and hurts paraphrase queries should not average to "no change".

D. Generator evaluation

The retriever finding the evidence is not the answer using it. Add:

Context Recall · Context Precision      (did retrieval cover the needed facts?)
Faithfulness / Groundedness             (is every claim supported by context?)
Answer Correctness / Completeness       (is the answer right and complete?)
Noise Sensitivity                       (does extra context hurt?)

RAGAS defines the context/generator split this way, and Microsoft's end-to-end RAG evaluation separates groundedness, relevance and completeness. Claim-level checking (RAGChecker-style atomic claims) fits this codebase better than answer-vs-reference embedding similarity.

E. Citation evaluation — P0 for this product

Citation is the differentiator, so it deserves first-class metrics rather than being implied by retrieval:

Citation Precision     cited spans that actually support their claim
Citation Recall        factual claims that carry a citation at all
Citation Entailment    claim ↔ evidence: cited ≠ correctly cited

This extends the coverage work already shipped in #156/#164. Page/block matching is deterministic; claim↔evidence entailment needs a judge. This is where KnowNote can be measurably better than a generic RAG stack, so it should not wait for a later epic.

F. Experiments

Ordered, cheapest-and-most-informative first. Fixed chunking within a sweep; never full Cartesian product.

1. candidateK sweep      dense and hybrid:  5 / 10 / 20 / 40     (contextK fixed)
2. contextK sweep        hybrid:            3 / 5 / 8           (candidateK fixed)
3. strategy ladder       dense → BM25 → hybrid RRF → hybrid + reranker
4. threshold sweep       0.3 / 0.4 / 0.5 / 0.6  (+ off)
5. chunking re-run       only after the above, on the new corpus

Every run records p50/p95 latency, index size and context tokens alongside quality. A quality gain that triples p95 is a product decision, not an automatic adoption — the existing rule already says this and should survive into v2.

G. Reporting

  • Dashboard of independent metrics, not a single scalar "KnowNote RAG score".
  • Every experiment reported as a delta against the production-config baseline.
  • Public-suite and in-domain numbers never averaged together.
  • Latency, index size and token cost are columns in the same table as the quality metrics.
                 Dense    Hybrid   Hybrid+Rerank
Recall@5          0.82      0.91       0.94
MRR               0.74      0.79       0.84
nDCG@10           0.77      0.83       0.88
Context Precision 0.56      0.62       0.79
Faithfulness      0.88      0.90       0.92
Citation Recall   0.76      0.84       0.90
p95               18 ms     34 ms     120 ms

H. Evaluation methodology

deterministic metrics  →  retrieval + citation location matching
LLM judge              →  faithfulness, completeness, entailment
small human audit      →  30-50 sampled cases per release

LLM-as-judge is non-deterministic and must not be the only evidence; a deterministic core plus a bounded human sample keeps release decisions reproducible.

Non-goals

  • No GraphRAG, Agentic RAG, HyDE, Self-RAG, ColBERT yet. Hybrid + RRF + reranker is the next real step for private-document QA with citations; anything more elaborate is premature until the benchmark can measure it.
  • No larger embedding model as the fix. Changing the model on a saturated 19-chunk corpus teaches nothing.
  • No single aggregate score.
  • No wholesale dataset replacement. The current first-party corpus stays as a smoke corpus; it is not deleted, it is demoted.

Known blocker

Reranking (#170) is blocked on a cross-encoder that can be pinned and run offline, the same way multilingual-e5-small@761b726 is pinned in eval:prepare. Until then, stage 1 is dense candidateK=20 + BM25 candidateK=20 → RRF → contextK=5, which needs no new model and is already a strictly better shape than the current slice(0, topK).

Acceptance criteria

  1. The harness runs the production configuration by default, and candidateK / contextK are separate named parameters in both the contract and the eval config.
  2. In-domain corpus is at least stage 1 scale, with hard-negative, unanswerable, multi-chunk and multi-document categories present and counted.
  3. Ground truth is quote/block-based and survives a re-chunk without re-labelling.
  4. Retrieval metrics are reported per query type; HitRate@K and MAP are added.
  5. Answer-level metrics (context precision/recall, faithfulness, completeness) exist and run on the same dataset.
  6. Citation precision / recall / entailment exist and are reported.
  7. The dense → hybrid → hybrid + reranker ladder is reported with candidateK and contextK sweeps, each with p50/p95, index size and context tokens.
  8. A public suite subset (at minimum QASPER + MIRACL-zh) runs reproducibly offline and is reported separately from in-domain numbers.
  9. docs/eval/ states, for the corpus in use, whether Recall@K is saturated — and the adoption rule is amended so saturation is detected rather than silently blocking every change.

Suggested child issues

# Issue Priority
1 Eval config parity: candidateK / contextK, production config as baseline P0
2 KnowNote Eval corpus v2 (scale, query taxonomy, graded ground truth, hard negatives, unanswerable) P0
3 Retrieval metrics v2: HitRate@K, MAP, per-query-type reporting P1
4 Re-derive threshold from the validation set; first stage threshold = off P1
5 Generator evaluation: context precision/recall, faithfulness, completeness P1
6 Citation evaluation: precision / recall / entailment P1
7 Parameter sweep harness + dashboard reporting (latency, index size, tokens) P1
8 Public suite runner: QASPER, MIRACL-zh (pinned offline subsets) P2
9 Pin an offline cross-encoder and evaluate hybrid + reranker P2 (blocked)
10 Amend the adoption rule to detect and report metric saturation P2

Build order

1 (config parity)  →  2 (corpus v2)  →  3 (metrics v2)
                                        ↓
                              4 (threshold)  →  7 (sweeps/dashboard)
                                        ↓
                              5 (generator)  →  6 (citation)
                                        ↓
                              8 (public)     →  9 (reranker)

Config parity and the corpus come first because every number produced before them is unrepresentative of the product.

References

  • LlamaIndex RetrieverEvaluator — hit rate, MRR, precision, recall, AP, nDCG
  • RAGAS — context precision, context recall, response relevancy, faithfulness, noise sensitivity
  • RAGChecker — atomic-claim precision/recall, claim recall, context utilization, hallucination
  • Datawhale All-in-RAG — hybrid retrieval, RRF, cross-encoder/ColBERT, context relevance & faithfulness
  • Microsoft — parameter sweep for algorithm/top-k/chunk-size; groundedness, relevance, completeness; LLM-judge non-determinism
  • NVIDIA RAG Blueprint — separate VDB top-k (100) and reranker top-k (10)
  • BEIR, MIRACL, QASPER, HAGRID, CRUD-RAG, RAGBench

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:provenanceSource location, citations, document structurearea:retrievalRetrieval quality and evaluationenhancementNew feature or requestepicTracking issue for a multi-PR epicpriority:P1High value, current or next milestoneresearch/spikeTimeboxed investigation; adopt only if the measurements support it

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions