Skip to content

Cite the most relevant sources, not the first collection's - #256

Merged
adamjohnwright merged 1 commit into
mainfrom
fix/rank-across-collections
Sep 19, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
fix/rank-across-collections

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

The contract tells the website to treat citations as "the most relevant few". Measured, that was false:

97 of 97 cited documents came from reactions.csv — including for "which diseases involve variants of the PTEN gene", where the disease_variants documents were retrieved and could never be cited.

Found while trying to measure whether collection routing would cost recall. It turned out the ranking that measurement depends on did not exist.

Cause 1: citations came from sub-retriever events

The hybrid retriever runs one BM25 and one vector search per query variant per collection — 51 on_retriever_end events for one question. astream_answer took citations from every one until it hit the cap, so it took them from whichever fired first: always the first collection's BM25. The fused result, which arrives last, was never used.

Sub-retriever events are now skipped, identified by the tags this repository sets itself ({collection}-bm25-{i}, {collection}-vector-{i}).

Not by class name — I checked, and the user guide's own chain-level retriever is a VectorStoreRetriever too. Filtering on the name would have silently dropped every user-guide citation, a feature added earlier today.

Cause 2: the fused result was not fused across collections

retrieve_documents ran reciprocal rank fusion within each collection and then concatenated, so the output was grouped — ten reactions, then ten complexes — rather than ranked.

That is issue #170 one level up. There, BM25 and vector results were joined end to end so the two were never scored against each other; the comment explaining it sits directly above the line that made the same mistake with collections. Both retrieval paths now fuse across collections, and the document count is unchanged, so this reorders rather than re-scopes.

After

question citations
Which diseases involve variants of PTEN 3 reactions, 3 complexes, 3 summations, 3 ewas, 3 disease_variants
Describe the components of the proteasome all five collections
How do I use the pathway browser? 4 from 4 user-guide pages — still working

Answer sweep 15/15, ruff, mypy (134 files), full suite with no API keys set.

Honest limit

RRF ranks by position, so with no overlap between collections the top of the list interleaves them rather than ordering by a comparable relevance score. That is a real improvement over grouping — every collection's best material now reaches the top and can be cited — but "most relevant few" is now approximately true rather than exactly so. Ranking across collections by a shared score would need rescoring, which is a larger change than this.

🤖 Generated with Claude Code

The contract tells the website to treat citations as "the most relevant few".
Measured, that was false: 97 of 97 cited documents came from reactions.csv,
including for "which diseases involve variants of the PTEN gene", where the
disease_variants documents were retrieved and could never be cited.

Two causes, both fixed.

**Citations came from sub-retriever events.** The hybrid retriever runs one BM25
and one vector search per query variant per collection, and each fires its own
on_retriever_end -- 51 events for one question. `astream_answer` took citations
from every one of them until it hit the cap, so it took them from whichever fired
first, which is always the first collection's BM25. The fused result, which
arrives last, was never used.

Sub-retriever events are now skipped, identified by the tags this repository sets
itself -- `{collection}-bm25-{i}` and `{collection}-vector-{i}`. Not by class
name: the user guide's own chain-level retriever is a VectorStoreRetriever too,
and filtering on that would have silently dropped every user-guide citation,
which is a feature added earlier today.

**And the fused result was not fused across collections.** `retrieve_documents`
ran reciprocal rank fusion within each collection and then concatenated, so the
output was grouped -- ten reactions, then ten complexes -- rather than ranked.
That is issue #170 one level up: there, BM25 and vector results were joined end
to end so the two were never scored against each other. Both retrieval paths now
fuse across collections, and the document count is unchanged, so this reorders
rather than re-scopes.

After: citations draw from all five collections, and disease_variants is cited
for the PTEN variants question. User-guide citations still work, 4 from 4 pages.
Answer sweep 15/15.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright

Copy link
Copy Markdown
Contributor Author

Adversarial review

Four checks. The one that mattered was whether reordering changed what the model sees, since the answer sweep is pattern-based and 15/15 is a weak signal on its own.

Did the fix change which documents reach the model? No — measured on two questions, before against after on genuinely different checkouts:

result
same document set True — 0 lost, 0 gained
same document order False

So the claim "reorders rather than re-scopes" is verified rather than asserted. The model sees the same evidence; only the order changed, which is what lets citations span collections.

Does the sync path still agree with the async one? tests/retrievers/test_sync_async_equivalence.py pins that they must agree document for document, and I changed both. It passes — and, mutation-checked, 3 of its tests fail when I break only the async path's fusion. So the invariant is genuinely enforced, not incidentally satisfied.

Is the tag-based filter safe? _is_sub_retriever keys on -bm25- and -vector-, which this repository sets itself in retrieve_documents. No collection directory name contains either substring (complexes, disease_variants, ewas, reactions, summations), so no chain-level event can be mistaken for a sub-retriever one.

Do all four answer paths still work? I edited astream_answer, which every path runs through:

path tokens citations state
reactome (vector store) 887 12 answered
userguide 428 4 answered
live MCP 54 0 answered
safety refusal 0 0 nothing_found

The userguide row is the one I most wanted to see: a name-based filter would have zeroed it, and it would have looked like the feature I added this morning had simply stopped working.

@adamjohnwright
adamjohnwright merged commit ea38aa4 into main Sep 19, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the fix/rank-across-collections branch September 19, 2026 01:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant