Skip to content

Answer-quality fixes (review, area 3) - #315

Merged
adamjohnwright merged 1 commit into
mainfrom
review-3-quality
Oct 4, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
review-3-quality

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

These come from the max-level review of the agent and retrieval code. Each one was reproduced by the reviewer.

Finding Fix Measured or tested
Disease-variant keyword search indexed all 21 columns, while the vectors embedded 12. RRF never fused the two, and results came back doubled. EMBEDDED_CONTENT_COLUMNS: keyword search indexes the embedded columns, limited to those present in the file On the real Release 97 bundle, 300 of 300 vector documents now match the keyword text (0 of 300 before)
Disease-variant citations had empty labels Labelled with the variant metadata Unit test
A failed grader or web search failed a turn that had already answered Logged and dropped Unit test
A safety-check refusal raised an error instead of giving the polite refusal OpenAIRefusalError is treated as an unsafe verdict Unit test
User-guide answers ignored the reader's language The user-guide prompt gets LANGUAGE_INSTRUCTION, in the same place as the Reactome prompt 3 runs each, old vs new prompt: French 0/3 → 3/3; Spanish 3/3 → 3/3; English 3/3 → 3/3. Answer sweep 16/16.
The classifier test compared the prompt with itself, so it could never fail Replaced: the no-MCP prompt must not mention live services Sabotaged: always including the live rule fails it

./checks.sh passes.

🤖 Generated with Claude Code

- Disease-variant keyword search indexes the 12 columns the vector store
  embedded, not all 21: the two retrievers never returned the same text
  for a variant, so RRF never fused them and results came back doubled.
  On the real bundle: 300 of 300 vector documents now match BM25 text
  (0 of 300 before).
- Disease-variant citations are labelled with their variant; they had no
  display_name and went out empty.
- A failed post-answer step (grader or web search) is logged and dropped;
  it failed a turn whose answer had already streamed and been saved.
- A model refusing the safety check's structured output is an unsafe
  verdict, so the reader gets the polite refusal, not an error.
- The user-guide prompt carries the language instruction. Measured, three
  runs each: French 0/3 -> 3/3 answered in French; Spanish and English
  unchanged. Answer sweep 16/16.
- The classifier test that compared a prompt with itself is replaced by
  one that fails when the no-MCP prompt mentions live services.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 75ca531 into main Oct 4, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the review-3-quality branch October 4, 2026 03:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant