Skip to content

Stage 2: plain semantic search replaces SelfQueryRetriever - #183

Merged
adamjohnwright merged 1 commit into
mainfrom
feat/retriever-stage2
Sep 9, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
feat/retriever-stage2

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

D1 — Helia's decision. Second stage of the plan.

What changed

The vector side is now vectordb.as_retriever(...) with no LLM in the loop.

Retrieval costs one LLM call per message instead of 21 — the query expansion. SelfQuery previously spent one call per collection per query variant, 4 × 5 = 20.

descriptions_info and field_info are removed from the retriever signature and its three rag.py call sites. Leaving required parameters nothing reads would keep callers coupled to metadata the retriever no longer uses.

metadata_info.py stays. evaluator.py and bin/retrieval_baseline still construct SelfQuery for comparison — deleting it would break the tool that measures this change.

Measured

Against the pre-rewrite retriever, 8 questions, fixed query set: all 8 changed. That is what the spec's 0.55 set overlap and 0.18 rank agreement predicted. This is a real change to what reaches the model; whether answers improve is the ragas work, not something this stage can claim.

Two predictions in the plan were wrong

"Stage 2 becomes deterministic once the LLM leaves the vector side." It did not — 5 of 8 questions still vary run to run. Traced it: the residual variance is in Chroma's approximate nearest-neighbour search. Plain vector search returns different results for the same query on 1 of 8 questions on reactions, with no LLM anywhere near it. Nothing removed here caused that and nothing here can fix it.

"Differences confined to the vector side." Every fused result changed, because the vector side feeds all of them. What is verifiable, and verified: BM25 is untouched — 8/8 identical for the same query twice, on both collections tested.

I would rather record both than quietly restate the criteria to match what happened.

Tests

Two, both checked for vacuity by reintroducing what they forbid:

  • the module must not import a self-query retriever again
  • RetrieverDict types its vector slot as BaseRetriever, not one implementation

118 tests, gates green.

🤖 Generated with Claude Code

D1, Helia's decision. The vector side is now vectordb.as_retriever(...) with no
LLM in the loop, so retrieval costs one LLM call per message -- the query
expansion -- instead of 21.

descriptions_info and field_info are removed from the retriever's signature
along with the three rag.py call sites that supplied them. Leaving required
parameters that nothing reads would keep callers coupled to metadata the
retriever no longer uses. metadata_info.py itself stays: the evaluator and
bin/retrieval_baseline still construct SelfQuery for comparison, and deleting it
would break the tool that measures this change.

Measured against the pre-rewrite retriever on 8 questions with a fixed query set:
all 8 changed, which is what the 0.55 set overlap and 0.18 rank agreement in the
spec predicted. This is a real change to what reaches the model, and answer
quality is the ragas work rather than something this stage can claim.

Two predictions in the plan were wrong, and the checks are more useful than the
predictions were:

  It said Stage 2 would become deterministic once the LLM left the vector side.
  It did not -- 5 of 8 questions still vary run to run. The residual variance is
  in Chroma's approximate nearest-neighbour search, measured directly: plain
  vector search returns different results for the same query on 1 of 8 questions
  for `reactions`. Nothing we removed caused it and nothing here can fix it.

  It said differences would be "confined to the vector side". Every fused result
  changed, because the vector side feeds all of them. What is verifiable, and
  verified, is that BM25 is untouched: 8/8 identical for the same query twice, on
  both collections tested.

Two tests pin the change: the module must not import a self-query retriever
again, and RetrieverDict types its vector slot as BaseRetriever rather than one
implementation. Both were checked for vacuity by reintroducing what they forbid.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit afb2b3e into main Sep 9, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the feat/retriever-stage2 branch September 9, 2026 13:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant