Repository navigation
Evaluate the rephrased question, as production does - #193
Merged
Merged
Conversation
The evaluator asked the RAG chain the raw question. generate_answer
passes `rephrased_input`, never the raw one, so retrieval in production
always happens on rewritten text.
That is not cosmetic. Measured over the 20 golden questions, 15 come
back changed:
signalling -> signaling
Describe the components of the proteasome
-> What are the components of the proteasome?
Which complexes contain histone deacetylases?
-> Which protein complexes include histone deacetylases?
BM25 is lexical, so a British-to-American spelling change alone
retrieves different documents. Skipping the rephrase left this file
measuring a different retrieval from the one it claims to measure --
the same defect as building a private retriever, one step further up,
and shipped in the commit that fixed the first one.
The model under test now rephrases its own questions, as production
does with a single LLM for both, and the report records the original
and the rephrasing side by side so a reader can see what was actually
retrieved on.
Scoring moved to the rephrased question too: judging an answer against
wording the product deliberately rewrote would penalise a faithful
answer for the rewrite.
That introduced a bug, fixed here rather than shipped: reference
answers are keyed by the question a person typed, so looking them up
with the rephrased text matched nothing. It would have dropped
context_recall from every run while looking exactly like "no references
were supplied". resolve_references() exists to make that testable
instead of asserting on source text.
Also documents what is NOT measured -- safety check, intent
classification, postprocessing -- and why none of them changes the
answer text.
adamjohnwright
deleted the
fix/evaluator-rephrases-as-production-does
branch
September 9, 2026 17:04
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The evaluator asked the RAG chain the raw question.
generate_answerpassesrephrased_input— never the raw one — so retrieval in production always happens on rewritten text.That is not cosmetic. Measured over the 20 golden questions, 15 come back changed:
BM25 is lexical, so the British→American spelling change alone retrieves different documents.
Skipping the rephrase left this file measuring a different retrieval from the one it claims to measure — the same defect as building a private retriever, one step further up, shipped in the commit that fixed the first one. Found by attacking my own change rather than by CI.
The model under test now rephrases its own questions, as production does with a single LLM for both, and the report records both:
Scoring moved to the rephrased question too — judging an answer against wording the product deliberately rewrote would penalise a faithful answer for the rewrite.
A bug that introduced, fixed here rather than shipped
Reference answers are keyed by the question a person typed. Looking them up with the rephrased text matches nothing — and would have dropped
context_recallfrom every run while looking exactly like "no references were supplied".resolve_references()was extracted so this is testable directly rather than by asserting on source text (a pattern that has produced false positives here twice). Perturbing it fails both new tests.Also documented
What is not measured, and why none of it changes the answer text: the safety check, intent classification (this always evaluates the reactome source), and postprocessing, which appends web-search results as separate content rather than rewriting the answer.