Repository navigation
Give BM25 documents the metadata the Chroma half already had - #247
Merged
Merged
Conversation
`create_bm25_chroma_ensemble_retriever` loaded its CSVs with plain CSVLoader, so
every BM25 document had metadata of only {row, source} -- while the Chroma half
of the same ensemble, built from the same CSV by MetaDataCSVLoader at embedding
time, had st_id and the rest.
The effect is on **citations**. A document with no `st_id` in its metadata cannot
be cited, even though its stable id sits in the page content as
"st_id: R-HSA-..." where nothing can reach it. Measured on "What does CDK5 do in
neurons?": of the documents the citation path walked, 50 of 62 were
unattributable before and 0 are now. So the sources a caller saw came only from
the Chroma half and under-represented what the answer drew on.
**What this does NOT do, despite my first claim: it does not shrink the model's
context.** I wrote that 62 retrieved documents became 14. That was a
mismeasurement -- the instrument hooked the citation loop, which stops at
MAX_CITATIONS, so it counted documents examined before the cap rather than
documents retrieved. Measured at `unique_documents`, where the model's context is
actually decided, and across three runs per question because retrieval is not
reproducible:
before: 50 docs in, 50 out, ~38 distinct stable ids
after: 50 docs in, 50 out, ~39 distinct stable ids
Unchanged. No duplicates are collapsed there in either version, so the "11 exact
duplicates" I reported were an artefact of the same instrument and not something
the model was being fed twice.
Content is untouched -- MetaDataCSVLoader trims content only when
`content_columns` is given, and it is not -- so BM25 scoring is unchanged, which
is consistent with retrieval being identical. Every column is promoted, read from
the header, so metadata cannot drift from the file.
The answer sweep stays green at 13/13.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nothing failed when they did not: the answers were still right, the sweep still passed, and the only symptom was citations quietly missing sources the answer had used. So the property is pinned, including that promoting columns does not trim the indexed text -- which is why retrieval measured identical before and after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Found while reviewing #246, which put a local filesystem path in front of me. The guard there stopped the leak; this removes the reason it happened.
The defect
create_bm25_chroma_ensemble_retrieverloads its CSVs with plainCSVLoader, so every BM25 document gets metadata of only{row, source}— while the Chroma half of the same ensemble, built from the same CSV byMetaDataCSVLoaderat embedding time, carriesst_idand the rest.A document with no
st_idin metadata cannot be cited, even though its stable id is sitting in the page content asst_id: R-HSA-...where nothing can reach it. On "What does CDK5 do in neurons?", of the documents the citation path walked, 50 of 62 were unattributable before and 0 are now. The sources a caller saw came only from the Chroma half.What this does NOT do — correcting my own first claim
My first version of this said 62 retrieved documents became 14, and that 11 exact duplicates were being fed to the model twice. Both were mismeasurements, caught by the adversarial pass on this PR:
MAX_CITATIONS— so it counted documents examined before the cap, not documents retrievedMeasured properly, at
unique_documentswhere the model's context is actually decided, three runs per question because retrieval is not reproducible:Unchanged. No duplicates are collapsed there in either version. The benefit of this change is citation completeness, and nothing else I can demonstrate.
That is consistent with the mechanism: content is untouched —
MetaDataCSVLoadertrims content only whencontent_columnsis given, and it is not — so BM25 scoring is identical and retrieval should not move.Verification
🤖 Generated with Claude Code