Skip to content

About

Reference implementation for RAG retrieval evaluation: BM25, dense and hybrid RRF search, cross-encoder reranking, bootstrap significance testing and failure analysis. Pluggable vector stores and model providers, tracing, 166 tests, CI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Agentic RAG Reliability Reference

CI License: MIT Python 3.11 | 3.12 Tests

A public reference for the retrieval and quality controls I use when designing production-oriented RAG systems. Generic data, no customer content, no private production logic. The point is the controls, not the corpus.

It runs end to end: python benchmark.py reproduces every number in reports/results.md.

The reasoning behind the implementation, eleven decisions with the rejected alternative and the cost of each, is in DESIGN.md.

What it demonstrates

  • stable Pydantic contracts for chunks and search results;
  • BM25 implemented directly, plus TF-IDF, so the lexical baseline is explainable rather than imported;
  • injected dense embeddings, compatible with local or hosted embedding services;
  • reciprocal-rank fusion for hybrid search, as a testable free function;
  • an optional cross-encoder reranker behind an interface, with latency measured rather than assumed;
  • offline metrics with Hit Rate@k and Recall@k kept apart, plus Precision@k, MRR and nDCG@k;
  • paired significance testing (McNemar and bootstrap intervals), because a delta on fifty questions is a claim about eight of them;
  • a failure taxonomy that separates ranking failures from retrieval failures, which is the difference between "build a reranker" and "a reranker cannot help you";
  • an annotation protocol with a self-agreement check;
  • deterministic tests, no external API calls.

Results on the bundled demo set

27 chunks, 33 labelled questions, intfloat/multilingual-e5-small. Full report in reports/results.md.

configuration hit@1 hit@5 recall@5 mrr@10
bm25 0.606 0.788 0.758 0.683
tfidf 0.606 0.788 0.758 0.690
dense 0.727 0.939 0.924 0.817
hybrid-rrf 0.667 0.879 0.864 0.760

Three findings worth more than the headline number:

The hybrid lost to pure dense. 0.667 against 0.727 at hit@1. Equal-weight RRF dilutes a strong leg with a weak one; fusion is not free. Weighting or query-dependent routing is the follow-up, not "add hybrid search" as a slogan.

The class breakdown contradicts the aggregate. Dense wins paraphrases by a wide margin (hit@1 0.500 vs 0.143), the two tie on exact keywords (0.900), and lexical wins factual questions (1.000 vs 0.889). Averaging these together produces one number that is true of no query class.

The improvement is not significant on hit@1. 0.606 → 0.727 looks decisive and is six discordant queries: five fixed, one broken. McNemar gives p = 0.219. On Recall@5 the paired bootstrap gives +0.167 [0.030, 0.318], which does exclude zero. Same experiment, two different honest conclusions depending on the metric, and neither is "we improved retrieval by 12 points".

Why the components are separated

The retriever should not know whether embeddings come from a provider API, a local sentence-transformer, or a self-hosted endpoint. The orchestrator should not be tied to one reranker. Evaluation should consume normalised hits, not vendor-specific responses. These boundaries are what make model changes and regression testing cheap.

flowchart LR
    Q[User query] --> L[Lexical: BM25 or TF-IDF]
    Q --> D[Dense retriever]
    L --> F[Reciprocal-rank fusion]
    D --> F
    F --> R[Optional reranker]
    R --> C[Context builder]
    C --> A[Answering agent]
    A --> T[Trace + eval record]
Loading

Run locally

python -m venv .venv
source .venv/bin/activate            # Windows: .venv\Scripts\activate
python -m pip install -r requirements-dev.txt
pytest -q                            # 166 tests, no network, no Docker, no keys

# Optional backends. Everything above stays green without them; installing
# them turns 7 skips into real coverage (in-process Qdrant, the Langfuse
# SDK-surface check). Only the four `integration` tests still want a server.
python -m pip install -r requirements-stores.txt -r requirements-providers.txt

# Reproduce the report. --embeddings hash needs no model; e5 downloads one.
python -m pip install -r requirements-dense.txt
python benchmark.py --embeddings e5 --report reports/results.md

Module map

module what it owns
schemas.py chunk / hit / result contracts
tokenize_ru.py Russian tokenisation; keeps negations, folds ё
bm25.py Okapi BM25, Lucene IDF variant
retrieval.py lexical, dense, hybrid retrievers; RRF
reranker.py cross-encoder adapter + latency measurement
evaluation.py Hit Rate, Recall, Precision, MRR, nDCG
statistics_tests.py McNemar, paired bootstrap, Wilson intervals
failure_analysis.py failure taxonomy, system-to-system diff
annotation.py golden-set validation, Cohen's kappa
benchmark.py runs every configuration, writes the report
stores/ one VectorStore contract over numpy, Qdrant and pgvector
providers/ ChatModel / EmbeddingModel over Ollama and GigaChat
observability/ tracing seam (null / Langfuse) and a dated price table

Two decisions worth knowing about

Two of the eleven in DESIGN.md, repeated here because they are the two most often got wrong.

Hit Rate@k and Recall@k are different metrics. Hit Rate asks whether any relevant chunk landed in the top k; Recall asks what share did. They coincide only when every question has exactly one relevant chunk, which is why conflating them is easy and invisible. An earlier version of evaluation.py computed Hit Rate under the name recall_at_k. The numbers were right for that set and the label was wrong; both now exist under their own names, and validate_golden_set reports the multi-relevant share so the assumption is stated rather than assumed.

The RRF constant is not universal. rrf_k defaults to 60, from Cormack et al. (2009). Qdrant's built-in fusion uses 2. Comparing a local implementation against a database's without pinning the constant on both sides produces two different rankings and a confusing afternoon.

Swappable backends, and what is actually proven

Three seams exist so the retrieval layer does not care what sits behind it. Each has its own document, including an explicit note on what the tests do not establish.

seam implementations document
vector store numpy (exact reference), Qdrant, pgvector docs/STORES.md
model provider Ollama, GigaChat docs/PROVIDERS.md
tracing and cost null (structured log), Langfuse docs/OBSERVABILITY.md
docker compose up -d                      # Qdrant + Postgres/pgvector
pytest tests/test_stores.py -q -rs        # green without Docker: integration tests skip

The store tests assert that identical vectors produce an identical top-k across all three backends, filters included: the choice of store does not move the metrics. Scope of that claim, stated so it is not overread: the parity set is small enough that both servers answer exactly, so it establishes the contract and the filter translation, not that approximate recall matches at scale. The measurement that would close the gap, ANN recall against brute force across ef_search values on a real corpus, is named in docs/STORES.md and has not been run here.

Production extensions

Deliberately out of scope here, and the list a production implementation needs:

  • persistent document and embedding versioning, chunk lineage, source timestamps;
  • tenant and permission filters applied before ranking, not after: filtering a top-k after the fact silently shortens results and leaks nothing but breaks recall;
  • query classification and adaptive retrieval depth;
  • context sufficiency checks and abstention rules;
  • golden sets versioned alongside prompts and models, with CI regression thresholds;
  • online feedback and drift monitoring, because an offline proxy is not quality.

Annotation

The labelling rule, the class taxonomy, the self-agreement check and the caveats that travel with any number from this set are in ANNOTATION_PROTOCOL.md.

About

Reference implementation for RAG retrieval evaluation: BM25, dense and hybrid RRF search, cross-encoder reranking, bootstrap significance testing and failure analysis. Pluggable vector stores and model providers, tracing, 166 tests, CI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages