Theme
One sentence: turn the retrieval regression test into a real RAG benchmark — one that runs the production configuration, measures the generator and the citations, and can actually discriminate between candidate algorithms.
The harness we built in v1.4/v1.5 is not the problem. dense + BM25 + RRF, Recall@K / MRR / nDCG@10, p50/p95 latency, a chunking A/B test, a pinned embedding model, and a reproducible offline run are already the hard part of the work. What we do not yet have is an evaluation whose numbers can be trusted to represent the product.
This epic is about fixing the conditions of the experiment, not replacing the algorithm.
Why the current numbers do not mean what they look like
1. The benchmark is saturated
eval/corpus/ is 13 documents / 19 chunks / 30 questions, evaluated at topK = 10. Retrieving 10 of 19 chunks per question means the harness hands the ranker ~53% of the entire knowledge base before scoring. docs/eval/retrieval-v1.5.md already says it plainly:
The rule's Recall@5 condition is saturated on this corpus: dense already scores 1.0000, so no strategy can improve it and the rule can therefore never be met here.
So the v1.5 result — hybrid better on Recall@1 (0.8667 vs 0.8333), MRR (0.9444 vs 0.9278) and nDCG@10 (0.9561 vs 0.9437) — is recorded as inconclusive not because hybrid is not better, but because the adoption rule ("Recall@5 must improve") cannot move on a corpus where Recall@5 is already perfect. Any future retrieval change is blocked by the same rule. That is a benchmark defect, not a retrieval result.
2. The harness does not run the product
|
Retrieval |
topK |
threshold |
Production (chatHandlers.ts) |
dense |
3 |
0.5 |
Eval (eval/run.ts) |
dense |
10 |
0 |
We are benchmarking a retriever the user never runs. A chunk at rank 4 with cosine 0.47 is a hit in the eval and does not exist in production. Recall@5 = 100% can be reported while the shipped path silently drops the evidence. This mismatch is more consequential than the value of topK itself.
3. Candidate K and context K are conflated
There is one topK in the contract. It is doing two different jobs:
- how wide the first stage searches (recall-oriented), and
- how much text is assembled into the prompt (precision/budget-oriented).
HybridRetriever.ts fuses and then slices at the same number:
hits = rrfFuse([denseHits, sparseHits]).slice(0, topK)
With topK = 3 in chat, dense contributes 3 and BM25 contributes 3 — the fusion pool is at most 6, then truncated back to 3. That is not a two-stage retrieval; it is a single narrow search with fusion applied inside it.
The standard shape separates these: retrieve a wide candidate pool, then narrow it.
┌→ Dense candidateK=20 ─┐
Query → Retrieval │ ├→ RRF → rerank → contextK=3..8 → LLM
└→ BM25 candidateK=20 ─┘
NVIDIA's RAG blueprint sets VDB top-k at 100 and reranker top-k at 10 for exactly this reason; the Datawhale advanced-RAG path retrieves top-20 then reranks to top-5.
4. Two things are simply not evaluated
docs/eval/retrieval-v1.5.md records that reranking is not evaluated (no offline cross-encoder could be pinned), and the harness stops at the retriever. There is no generator metric and no citation metric — while citation is the product's core differentiator and has shipped since #155/#156/#162/#164.
Scope
A. Evaluation dataset
Two suites, deliberately separate.
A1. In-domain — KnowNote Eval (the one that sets production defaults)
Public datasets test their own distribution, not the one users upload. Target scale, staged:
stage 1: ~20 documents / ~300 chunks / ~80 questions
stage 2: 50-100 docs / 500-3000 chunks / 200-500 questions
Source types: PDF, Markdown, Web, DOCX, PPTX. Query types deliberately covered, because these are what separate algorithms:
exact keyword · semantic · paraphrase · single-hop · multi-chunk
multi-document · entity/number/version · hard negative · unanswerable
source-scoped · citation-required
Hard negatives are mandatory — e.g. two chunks for "GPT-4 context window" and "GPT-4 Turbo context window" where dense cross-contaminates. Unanswerable questions are mandatory: today the harness has no way to measure whether we say "not in your sources".
Ground truth extends the existing relevant[] schema with graded relevance, a reference answer, and claims, so nDCG can use graded gains and the generator/citation layers have a target:
{
"id": "kn-001",
"question": "...",
"answerable": true,
"type": "multi-source",
"referenceAnswer": "...",
"claims": ["...", "..."],
"evidence": [
{ "document": "...", "page": 12, "block": 21, "quote": "...", "relevance": 3 }
]
}
relevance: 3 essential / 2 useful / 1 supporting / 0 irrelevant.
Ground truth must stay quote-and-location based (block/page), not chunk-id based — otherwise re-chunking invalidates the dataset and chunk-size experiments become unmeasurable. The chunk↔block mapping from #108 is what makes this work.
LLM-authored questions are allowed as candidates but must be reviewed, and the hard categories (paraphrase, hard negative, multi-hop, multi-source, unanswerable) authored deliberately. A fully synthetic set reproduces the current failure: the source sentence and the generated question share vocabulary, dense scores 1.0, and we are back to a saturated benchmark.
A2. Public — sanity and comparability
Kept small and pinned; a reproducible subset is enough. Priority order:
| Dataset |
Tests |
Why |
| QASPER |
long-doc QA + supporting evidence |
closest to "upload a PDF, ask, cite the paragraph"; 55.5% of questions need multiple paragraphs |
| MIRACL-zh |
multilingual retrieval |
official human-judged zh subset; exercises multilingual-e5-small vs BM25 vs hybrid |
| BEIR / SciFact |
standard IR |
external sanity check that our retriever is a normal retriever |
| HAGRID |
answer + attribution |
reference for citation evaluation |
| CRUD-RAG |
Chinese end-to-end RAG |
zh generator-side reference |
| RAGBench |
end-to-end RAG |
relevance / adherence / completeness / utilization |
The public suite answers "is our retriever competitive?" The in-domain suite answers "did the product get better?" Only the second one sets defaults. Public results must always be reported separately and labelled as out-of-distribution.
B. Config parity — one config, both paths
- Split the contract:
candidateK (first stage, per channel) and contextK (assembled into the prompt). Keep threshold explicit as a separate concern.
- The eval harness must run the production configuration as its baseline run, and any deviation must be a named experiment.
HybridRetriever must fuse a wide pool and truncate once, after fusion — not slice each channel to the final K.
- Re-derive
threshold from the validation set instead of hand-picking 0.5. Cosine ~0.5 has no universal meaning; it varies with model, language, query type, chunk length and domain. Sweep 0.3 / 0.4 / 0.5 / 0.6 and watch recall, precision and the no-result rate. First stage may legitimately run threshold = off and let a reranker do the filtering.
C. Retrieval metrics
Keep Recall@K, MRR, nDCG@K. Add HitRate@K, Precision@K and MAP. Report per query type, not only aggregate — a strategy that helps entity queries and hurts paraphrase queries should not average to "no change".
D. Generator evaluation
The retriever finding the evidence is not the answer using it. Add:
Context Recall · Context Precision (did retrieval cover the needed facts?)
Faithfulness / Groundedness (is every claim supported by context?)
Answer Correctness / Completeness (is the answer right and complete?)
Noise Sensitivity (does extra context hurt?)
RAGAS defines the context/generator split this way, and Microsoft's end-to-end RAG evaluation separates groundedness, relevance and completeness. Claim-level checking (RAGChecker-style atomic claims) fits this codebase better than answer-vs-reference embedding similarity.
E. Citation evaluation — P0 for this product
Citation is the differentiator, so it deserves first-class metrics rather than being implied by retrieval:
Citation Precision cited spans that actually support their claim
Citation Recall factual claims that carry a citation at all
Citation Entailment claim ↔ evidence: cited ≠ correctly cited
This extends the coverage work already shipped in #156/#164. Page/block matching is deterministic; claim↔evidence entailment needs a judge. This is where KnowNote can be measurably better than a generic RAG stack, so it should not wait for a later epic.
F. Experiments
Ordered, cheapest-and-most-informative first. Fixed chunking within a sweep; never full Cartesian product.
1. candidateK sweep dense and hybrid: 5 / 10 / 20 / 40 (contextK fixed)
2. contextK sweep hybrid: 3 / 5 / 8 (candidateK fixed)
3. strategy ladder dense → BM25 → hybrid RRF → hybrid + reranker
4. threshold sweep 0.3 / 0.4 / 0.5 / 0.6 (+ off)
5. chunking re-run only after the above, on the new corpus
Every run records p50/p95 latency, index size and context tokens alongside quality. A quality gain that triples p95 is a product decision, not an automatic adoption — the existing rule already says this and should survive into v2.
G. Reporting
- Dashboard of independent metrics, not a single scalar "KnowNote RAG score".
- Every experiment reported as a delta against the production-config baseline.
- Public-suite and in-domain numbers never averaged together.
- Latency, index size and token cost are columns in the same table as the quality metrics.
Dense Hybrid Hybrid+Rerank
Recall@5 0.82 0.91 0.94
MRR 0.74 0.79 0.84
nDCG@10 0.77 0.83 0.88
Context Precision 0.56 0.62 0.79
Faithfulness 0.88 0.90 0.92
Citation Recall 0.76 0.84 0.90
p95 18 ms 34 ms 120 ms
H. Evaluation methodology
deterministic metrics → retrieval + citation location matching
LLM judge → faithfulness, completeness, entailment
small human audit → 30-50 sampled cases per release
LLM-as-judge is non-deterministic and must not be the only evidence; a deterministic core plus a bounded human sample keeps release decisions reproducible.
Non-goals
- No GraphRAG, Agentic RAG, HyDE, Self-RAG, ColBERT yet. Hybrid + RRF + reranker is the next real step for private-document QA with citations; anything more elaborate is premature until the benchmark can measure it.
- No larger embedding model as the fix. Changing the model on a saturated 19-chunk corpus teaches nothing.
- No single aggregate score.
- No wholesale dataset replacement. The current first-party corpus stays as a smoke corpus; it is not deleted, it is demoted.
Known blocker
Reranking (#170) is blocked on a cross-encoder that can be pinned and run offline, the same way multilingual-e5-small@761b726 is pinned in eval:prepare. Until then, stage 1 is dense candidateK=20 + BM25 candidateK=20 → RRF → contextK=5, which needs no new model and is already a strictly better shape than the current slice(0, topK).
Acceptance criteria
- The harness runs the production configuration by default, and
candidateK / contextK are separate named parameters in both the contract and the eval config.
- In-domain corpus is at least stage 1 scale, with hard-negative, unanswerable, multi-chunk and multi-document categories present and counted.
- Ground truth is quote/block-based and survives a re-chunk without re-labelling.
- Retrieval metrics are reported per query type; HitRate@K and MAP are added.
- Answer-level metrics (context precision/recall, faithfulness, completeness) exist and run on the same dataset.
- Citation precision / recall / entailment exist and are reported.
- The
dense → hybrid → hybrid + reranker ladder is reported with candidateK and contextK sweeps, each with p50/p95, index size and context tokens.
- A public suite subset (at minimum QASPER + MIRACL-zh) runs reproducibly offline and is reported separately from in-domain numbers.
docs/eval/ states, for the corpus in use, whether Recall@K is saturated — and the adoption rule is amended so saturation is detected rather than silently blocking every change.
Suggested child issues
| # |
Issue |
Priority |
| 1 |
Eval config parity: candidateK / contextK, production config as baseline |
P0 |
| 2 |
KnowNote Eval corpus v2 (scale, query taxonomy, graded ground truth, hard negatives, unanswerable) |
P0 |
| 3 |
Retrieval metrics v2: HitRate@K, MAP, per-query-type reporting |
P1 |
| 4 |
Re-derive threshold from the validation set; first stage threshold = off |
P1 |
| 5 |
Generator evaluation: context precision/recall, faithfulness, completeness |
P1 |
| 6 |
Citation evaluation: precision / recall / entailment |
P1 |
| 7 |
Parameter sweep harness + dashboard reporting (latency, index size, tokens) |
P1 |
| 8 |
Public suite runner: QASPER, MIRACL-zh (pinned offline subsets) |
P2 |
| 9 |
Pin an offline cross-encoder and evaluate hybrid + reranker |
P2 (blocked) |
| 10 |
Amend the adoption rule to detect and report metric saturation |
P2 |
Build order
1 (config parity) → 2 (corpus v2) → 3 (metrics v2)
↓
4 (threshold) → 7 (sweeps/dashboard)
↓
5 (generator) → 6 (citation)
↓
8 (public) → 9 (reranker)
Config parity and the corpus come first because every number produced before them is unrepresentative of the product.
References
- LlamaIndex
RetrieverEvaluator — hit rate, MRR, precision, recall, AP, nDCG
- RAGAS — context precision, context recall, response relevancy, faithfulness, noise sensitivity
- RAGChecker — atomic-claim precision/recall, claim recall, context utilization, hallucination
- Datawhale All-in-RAG — hybrid retrieval, RRF, cross-encoder/ColBERT, context relevance & faithfulness
- Microsoft — parameter sweep for algorithm/top-k/chunk-size; groundedness, relevance, completeness; LLM-judge non-determinism
- NVIDIA RAG Blueprint — separate VDB top-k (100) and reranker top-k (10)
- BEIR, MIRACL, QASPER, HAGRID, CRUD-RAG, RAGBench
Theme
One sentence: turn the retrieval regression test into a real RAG benchmark — one that runs the production configuration, measures the generator and the citations, and can actually discriminate between candidate algorithms.
The harness we built in v1.4/v1.5 is not the problem.
dense + BM25 + RRF, Recall@K / MRR / nDCG@10, p50/p95 latency, a chunking A/B test, a pinned embedding model, and a reproducible offline run are already the hard part of the work. What we do not yet have is an evaluation whose numbers can be trusted to represent the product.This epic is about fixing the conditions of the experiment, not replacing the algorithm.
Why the current numbers do not mean what they look like
1. The benchmark is saturated
eval/corpus/is 13 documents / 19 chunks / 30 questions, evaluated attopK = 10. Retrieving 10 of 19 chunks per question means the harness hands the ranker ~53% of the entire knowledge base before scoring.docs/eval/retrieval-v1.5.mdalready says it plainly:So the v1.5 result — hybrid better on Recall@1 (0.8667 vs 0.8333), MRR (0.9444 vs 0.9278) and nDCG@10 (0.9561 vs 0.9437) — is recorded as inconclusive not because hybrid is not better, but because the adoption rule ("Recall@5 must improve") cannot move on a corpus where Recall@5 is already perfect. Any future retrieval change is blocked by the same rule. That is a benchmark defect, not a retrieval result.
2. The harness does not run the product
chatHandlers.ts)eval/run.ts)We are benchmarking a retriever the user never runs. A chunk at rank 4 with cosine 0.47 is a hit in the eval and does not exist in production.
Recall@5 = 100%can be reported while the shipped path silently drops the evidence. This mismatch is more consequential than the value oftopKitself.3. Candidate K and context K are conflated
There is one
topKin the contract. It is doing two different jobs:HybridRetriever.tsfuses and then slices at the same number:With
topK = 3in chat, dense contributes 3 and BM25 contributes 3 — the fusion pool is at most 6, then truncated back to 3. That is not a two-stage retrieval; it is a single narrow search with fusion applied inside it.The standard shape separates these: retrieve a wide candidate pool, then narrow it.
NVIDIA's RAG blueprint sets VDB top-k at 100 and reranker top-k at 10 for exactly this reason; the Datawhale advanced-RAG path retrieves top-20 then reranks to top-5.
4. Two things are simply not evaluated
docs/eval/retrieval-v1.5.mdrecords that reranking is not evaluated (no offline cross-encoder could be pinned), and the harness stops at the retriever. There is no generator metric and no citation metric — while citation is the product's core differentiator and has shipped since #155/#156/#162/#164.Scope
A. Evaluation dataset
Two suites, deliberately separate.
A1. In-domain —
KnowNote Eval(the one that sets production defaults)Public datasets test their own distribution, not the one users upload. Target scale, staged:
Source types: PDF, Markdown, Web, DOCX, PPTX. Query types deliberately covered, because these are what separate algorithms:
Hard negatives are mandatory — e.g. two chunks for "GPT-4 context window" and "GPT-4 Turbo context window" where dense cross-contaminates. Unanswerable questions are mandatory: today the harness has no way to measure whether we say "not in your sources".
Ground truth extends the existing
relevant[]schema with graded relevance, a reference answer, and claims, so nDCG can use graded gains and the generator/citation layers have a target:{ "id": "kn-001", "question": "...", "answerable": true, "type": "multi-source", "referenceAnswer": "...", "claims": ["...", "..."], "evidence": [ { "document": "...", "page": 12, "block": 21, "quote": "...", "relevance": 3 } ] }relevance: 3 essential / 2 useful / 1 supporting / 0 irrelevant.Ground truth must stay quote-and-location based (block/page), not chunk-id based — otherwise re-chunking invalidates the dataset and chunk-size experiments become unmeasurable. The chunk↔block mapping from #108 is what makes this work.
LLM-authored questions are allowed as candidates but must be reviewed, and the hard categories (paraphrase, hard negative, multi-hop, multi-source, unanswerable) authored deliberately. A fully synthetic set reproduces the current failure: the source sentence and the generated question share vocabulary, dense scores 1.0, and we are back to a saturated benchmark.
A2. Public — sanity and comparability
Kept small and pinned; a reproducible subset is enough. Priority order:
multilingual-e5-smallvs BM25 vs hybridThe public suite answers "is our retriever competitive?" The in-domain suite answers "did the product get better?" Only the second one sets defaults. Public results must always be reported separately and labelled as out-of-distribution.
B. Config parity — one config, both paths
candidateK(first stage, per channel) andcontextK(assembled into the prompt). Keepthresholdexplicit as a separate concern.HybridRetrievermust fuse a wide pool and truncate once, after fusion — not slice each channel to the final K.thresholdfrom the validation set instead of hand-picking 0.5. Cosine ~0.5 has no universal meaning; it varies with model, language, query type, chunk length and domain. Sweep0.3 / 0.4 / 0.5 / 0.6and watch recall, precision and the no-result rate. First stage may legitimately runthreshold = offand let a reranker do the filtering.C. Retrieval metrics
Keep Recall@K, MRR, nDCG@K. Add HitRate@K, Precision@K and MAP. Report per query type, not only aggregate — a strategy that helps entity queries and hurts paraphrase queries should not average to "no change".
D. Generator evaluation
The retriever finding the evidence is not the answer using it. Add:
RAGAS defines the context/generator split this way, and Microsoft's end-to-end RAG evaluation separates groundedness, relevance and completeness. Claim-level checking (RAGChecker-style atomic claims) fits this codebase better than answer-vs-reference embedding similarity.
E. Citation evaluation — P0 for this product
Citation is the differentiator, so it deserves first-class metrics rather than being implied by retrieval:
This extends the coverage work already shipped in #156/#164. Page/block matching is deterministic; claim↔evidence entailment needs a judge. This is where KnowNote can be measurably better than a generic RAG stack, so it should not wait for a later epic.
F. Experiments
Ordered, cheapest-and-most-informative first. Fixed chunking within a sweep; never full Cartesian product.
Every run records p50/p95 latency, index size and context tokens alongside quality. A quality gain that triples p95 is a product decision, not an automatic adoption — the existing rule already says this and should survive into v2.
G. Reporting
H. Evaluation methodology
LLM-as-judge is non-deterministic and must not be the only evidence; a deterministic core plus a bounded human sample keeps release decisions reproducible.
Non-goals
Known blocker
Reranking (#170) is blocked on a cross-encoder that can be pinned and run offline, the same way
multilingual-e5-small@761b726is pinned ineval:prepare. Until then, stage 1 isdense candidateK=20 + BM25 candidateK=20 → RRF → contextK=5, which needs no new model and is already a strictly better shape than the currentslice(0, topK).Acceptance criteria
candidateK/contextKare separate named parameters in both the contract and the eval config.dense → hybrid → hybrid + rerankerladder is reported with candidateK and contextK sweeps, each with p50/p95, index size and context tokens.docs/eval/states, for the corpus in use, whether Recall@K is saturated — and the adoption rule is amended so saturation is detected rather than silently blocking every change.Suggested child issues
candidateK/contextK, production config as baselinethresholdfrom the validation set; first stagethreshold = offBuild order
Config parity and the corpus come first because every number produced before them is unrepresentative of the product.
References
RetrieverEvaluator— hit rate, MRR, precision, recall, AP, nDCG