Skip to content

Give BM25 documents the metadata the Chroma half already had - #247

Merged
adamjohnwright merged 2 commits into
mainfrom
fix/bm25-document-metadata
Sep 18, 2026
Merged

adamjohnwright merged 2 commits into
mainfrom
fix/bm25-document-metadata

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Found while reviewing #246, which put a local filesystem path in front of me. The guard there stopped the leak; this removes the reason it happened.

The defect

create_bm25_chroma_ensemble_retriever loads its CSVs with plain CSVLoader, so every BM25 document gets metadata of only {row, source} — while the Chroma half of the same ensemble, built from the same CSV by MetaDataCSVLoader at embedding time, carries st_id and the rest.

A document with no st_id in metadata cannot be cited, even though its stable id is sitting in the page content as st_id: R-HSA-... where nothing can reach it. On "What does CDK5 do in neurons?", of the documents the citation path walked, 50 of 62 were unattributable before and 0 are now. The sources a caller saw came only from the Chroma half.

What this does NOT do — correcting my own first claim

My first version of this said 62 retrieved documents became 14, and that 11 exact duplicates were being fed to the model twice. Both were mismeasurements, caught by the adversarial pass on this PR:

  • the instrument hooked the citation loop, which stops at MAX_CITATIONS — so it counted documents examined before the cap, not documents retrieved
  • and my "before" run had nothing to stash, because I had already committed, so it measured the fixed code against itself

Measured properly, at unique_documents where the model's context is actually decided, three runs per question because retrieval is not reproducible:

docs into dedup docs to model distinct st_ids
before (main) 50 50 ~38
after (fix) 50 50 ~39

Unchanged. No duplicates are collapsed there in either version. The benefit of this change is citation completeness, and nothing else I can demonstrate.

That is consistent with the mechanism: content is untouched — MetaDataCSVLoader trims content only when content_columns is given, and it is not — so BM25 scoring is identical and retrieval should not move.

Verification

  • answer sweep green, 13/13 (2 skipped for no local MCP, as expected)
  • every column promoted, read from the header, so metadata cannot drift from the file
  • ruff, format, mypy (130 files), full suite with no API keys set

🤖 Generated with Claude Code

adamjohnwright and others added 2 commits September 18, 2026 06:38
`create_bm25_chroma_ensemble_retriever` loaded its CSVs with plain CSVLoader, so
every BM25 document had metadata of only {row, source} -- while the Chroma half
of the same ensemble, built from the same CSV by MetaDataCSVLoader at embedding
time, had st_id and the rest.

The effect is on **citations**. A document with no `st_id` in its metadata cannot
be cited, even though its stable id sits in the page content as
"st_id: R-HSA-..." where nothing can reach it. Measured on "What does CDK5 do in
neurons?": of the documents the citation path walked, 50 of 62 were
unattributable before and 0 are now. So the sources a caller saw came only from
the Chroma half and under-represented what the answer drew on.

**What this does NOT do, despite my first claim: it does not shrink the model's
context.** I wrote that 62 retrieved documents became 14. That was a
mismeasurement -- the instrument hooked the citation loop, which stops at
MAX_CITATIONS, so it counted documents examined before the cap rather than
documents retrieved. Measured at `unique_documents`, where the model's context is
actually decided, and across three runs per question because retrieval is not
reproducible:

    before: 50 docs in, 50 out, ~38 distinct stable ids
    after:  50 docs in, 50 out, ~39 distinct stable ids

Unchanged. No duplicates are collapsed there in either version, so the "11 exact
duplicates" I reported were an artefact of the same instrument and not something
the model was being fed twice.

Content is untouched -- MetaDataCSVLoader trims content only when
`content_columns` is given, and it is not -- so BM25 scoring is unchanged, which
is consistent with retrieval being identical. Every column is promoted, read from
the header, so metadata cannot drift from the file.

The answer sweep stays green at 13/13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nothing failed when they did not: the answers were still right, the sweep still
passed, and the only symptom was citations quietly missing sources the answer had
used. So the property is pinned, including that promoting columns does not trim
the indexed text -- which is why retrieval measured identical before and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 0732589 into main Sep 18, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the fix/bm25-document-metadata branch September 18, 2026 06:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant