Cite the most relevant sources, not the first collection's - #256
Conversation
The contract tells the website to treat citations as "the most relevant few".
Measured, that was false: 97 of 97 cited documents came from reactions.csv,
including for "which diseases involve variants of the PTEN gene", where the
disease_variants documents were retrieved and could never be cited.
Two causes, both fixed.
**Citations came from sub-retriever events.** The hybrid retriever runs one BM25
and one vector search per query variant per collection, and each fires its own
on_retriever_end -- 51 events for one question. `astream_answer` took citations
from every one of them until it hit the cap, so it took them from whichever fired
first, which is always the first collection's BM25. The fused result, which
arrives last, was never used.
Sub-retriever events are now skipped, identified by the tags this repository sets
itself -- `{collection}-bm25-{i}` and `{collection}-vector-{i}`. Not by class
name: the user guide's own chain-level retriever is a VectorStoreRetriever too,
and filtering on that would have silently dropped every user-guide citation,
which is a feature added earlier today.
**And the fused result was not fused across collections.** `retrieve_documents`
ran reciprocal rank fusion within each collection and then concatenated, so the
output was grouped -- ten reactions, then ten complexes -- rather than ranked.
That is issue #170 one level up: there, BM25 and vector results were joined end
to end so the two were never scored against each other. Both retrieval paths now
fuse across collections, and the document count is unchanged, so this reorders
rather than re-scopes.
After: citations draw from all five collections, and disease_variants is cited
for the PTEN variants question. User-guide citations still work, 4 from 4 pages.
Answer sweep 15/15.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adversarial reviewFour checks. The one that mattered was whether reordering changed what the model sees, since the answer sweep is pattern-based and 15/15 is a weak signal on its own. Did the fix change which documents reach the model? No — measured on two questions, before against after on genuinely different checkouts:
So the claim "reorders rather than re-scopes" is verified rather than asserted. The model sees the same evidence; only the order changed, which is what lets citations span collections. Does the sync path still agree with the async one? Is the tag-based filter safe? Do all four answer paths still work? I edited
The userguide row is the one I most wanted to see: a name-based filter would have zeroed it, and it would have looked like the feature I added this morning had simply stopped working. |
The contract tells the website to treat citations as "the most relevant few". Measured, that was false:
Found while trying to measure whether collection routing would cost recall. It turned out the ranking that measurement depends on did not exist.
Cause 1: citations came from sub-retriever events
The hybrid retriever runs one BM25 and one vector search per query variant per collection — 51
on_retriever_endevents for one question.astream_answertook citations from every one until it hit the cap, so it took them from whichever fired first: always the first collection's BM25. The fused result, which arrives last, was never used.Sub-retriever events are now skipped, identified by the tags this repository sets itself (
{collection}-bm25-{i},{collection}-vector-{i}).Not by class name — I checked, and the user guide's own chain-level retriever is a
VectorStoreRetrievertoo. Filtering on the name would have silently dropped every user-guide citation, a feature added earlier today.Cause 2: the fused result was not fused across collections
retrieve_documentsran reciprocal rank fusion within each collection and then concatenated, so the output was grouped — ten reactions, then ten complexes — rather than ranked.That is issue #170 one level up. There, BM25 and vector results were joined end to end so the two were never scored against each other; the comment explaining it sits directly above the line that made the same mistake with collections. Both retrieval paths now fuse across collections, and the document count is unchanged, so this reorders rather than re-scopes.
After
Answer sweep 15/15, ruff, mypy (134 files), full suite with no API keys set.
Honest limit
RRF ranks by position, so with no overlap between collections the top of the list interleaves them rather than ordering by a comparable relevance score. That is a real improvement over grouping — every collection's best material now reaches the top and can be cited — but "most relevant few" is now approximately true rather than exactly so. Ranking across collections by a shared score would need rescoring, which is a larger change than this.
🤖 Generated with Claude Code