Skip to content

Evaluate the rephrased question, as production does - #193

Merged
adamjohnwright merged 1 commit into
mainfrom
fix/evaluator-rephrases-as-production-does
Sep 9, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
fix/evaluator-rephrases-as-production-does

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

The evaluator asked the RAG chain the raw question. generate_answer passes rephrased_input — never the raw one — so retrieval in production always happens on rewritten text.

That is not cosmetic. Measured over the 20 golden questions, 15 come back changed:

signalling                                    -> signaling
Describe the components of the proteasome     -> What are the components of the proteasome?
Which complexes contain histone deacetylases? -> Which protein complexes include histone deacetylases?

BM25 is lexical, so the British→American spelling change alone retrieves different documents.

Skipping the rephrase left this file measuring a different retrieval from the one it claims to measure — the same defect as building a private retriever, one step further up, shipped in the commit that fixed the first one. Found by attacking my own change rather than by CI.

The model under test now rephrases its own questions, as production does with a single LLM for both, and the report records both:

question : What does CDK5 phosphorylate in Alzheimer's disease?
rephrased: What are the substrates that CDK5 phosphorylates in the context of Alzheimer's disease?

Scoring moved to the rephrased question too — judging an answer against wording the product deliberately rewrote would penalise a faithful answer for the rewrite.

A bug that introduced, fixed here rather than shipped

Reference answers are keyed by the question a person typed. Looking them up with the rephrased text matches nothing — and would have dropped context_recall from every run while looking exactly like "no references were supplied".

resolve_references() was extracted so this is testable directly rather than by asserting on source text (a pattern that has produced false positives here twice). Perturbing it fails both new tests.

Also documented

What is not measured, and why none of it changes the answer text: the safety check, intent classification (this always evaluates the reactome source), and postprocessing, which appends web-search results as separate content rather than rewriting the answer.

The evaluator asked the RAG chain the raw question. generate_answer
passes `rephrased_input`, never the raw one, so retrieval in production
always happens on rewritten text.

That is not cosmetic. Measured over the 20 golden questions, 15 come
back changed:

    signalling -> signaling
    Describe the components of the proteasome
      -> What are the components of the proteasome?
    Which complexes contain histone deacetylases?
      -> Which protein complexes include histone deacetylases?

BM25 is lexical, so a British-to-American spelling change alone
retrieves different documents. Skipping the rephrase left this file
measuring a different retrieval from the one it claims to measure --
the same defect as building a private retriever, one step further up,
and shipped in the commit that fixed the first one.

The model under test now rephrases its own questions, as production
does with a single LLM for both, and the report records the original
and the rephrasing side by side so a reader can see what was actually
retrieved on.

Scoring moved to the rephrased question too: judging an answer against
wording the product deliberately rewrote would penalise a faithful
answer for the rewrite.

That introduced a bug, fixed here rather than shipped: reference
answers are keyed by the question a person typed, so looking them up
with the rephrased text matched nothing. It would have dropped
context_recall from every run while looking exactly like "no references
were supplied". resolve_references() exists to make that testable
instead of asserting on source text.

Also documents what is NOT measured -- safety check, intent
classification, postprocessing -- and why none of them changes the
answer text.
@adamjohnwright
adamjohnwright merged commit b88fd02 into main Sep 9, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the fix/evaluator-rephrases-as-production-does branch September 9, 2026 17:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant