Skip to content

Commit 96c5145

Browse files
Close out spec 004: implemented, #140 and #125 closed with credit
T012-T014. All 15 tasks complete. Each contributor was told specifically what of theirs was used. #140's instruction wording is in language_instruction.py nearly verbatim, because it got the part that is easy to miss: nomenclature must survive untranslated, and SET, MAX and CAT are gene symbols as well as English words. #125's mechanism -- a prompt variable rather than a concatenated query -- is the one #205 uses. #140 was also told where the 0-of-10 figure came from and that it was wrong, since a dramatic number in a rejection deserves an explanation. #125 was told plainly that closing it is not a decision about its hallucination grader, which belongs to #123 and is worth having. Records both wrong claims and the lesson behind the second: the first review checked whether each claim was true without checking whether the test measured the product, which is the same failure evaluator.py had four days earlier. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent 0454beb commit 96c5145

2 files changed

Lines changed: 60 additions & 4 deletions

File tree

‎specs/004-answer-in-user-language/spec.md‎

Lines changed: 57 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,9 @@
44

55
**Created**: 2026-09-10
66

7-
**Status**: Draft. Two contributed PRs to harvest; one decision (D1) for the team.
7+
**Status**: **Implemented** 2026-09-10 (#204 spec, #205 implementation). D1 taken as
8+
recommended — instruct the answer prompt, no translation pass. #140 and #125 closed
9+
with credit. Outcome at the end of this file.
810

911
**Input**: Issue #104, "RAG only responds in English". Two competing pull requests, #125 and #140.
1012

@@ -297,3 +299,57 @@ which is what both contributors did, independently.
297299
- The web-search and hallucination-grading work bundled into #125 — that is #123's
298300
decision to make.
299301
- The user interface: language is detected per message, not chosen in a setting.
302+
303+
304+
---
305+
306+
## Outcome (2026-09-10)
307+
308+
Implemented in #205. `detected_language` reaches the answer prompt in React-to-Me and
309+
Plant Reactome as its own variable; `input` is untouched, so the retrieval query and
310+
the query expansion in front of it are byte-identical across languages. A test
311+
asserts that directly, which is exact where a retrieval baseline would only show
312+
noise.
313+
314+
Verified against the Release95 bundle: a French question is answered in French with
315+
`R-HSA` identifiers intact, an English question is unchanged, both retrieve 40
316+
documents.
317+
318+
### What the reviews changed
319+
320+
The specification's own claims were wrong twice, both caught before implementation.
321+
322+
**"BM25 retrieves close to nothing for a French query"** — false. It returns a full
323+
ten documents; named entities survive translation. The real cost is that seven of ten
324+
differ.
325+
326+
**"#140's appended instruction leaves 0 of 10 documents"** — measured on BM25
327+
directly, which the pipeline never does: the query expander rewrites the question
328+
first, so four of five queries reach BM25 clean. Through the whole retriever it is
329+
**20 of 40**. Half the context, not all of it. Adam pushed back that the number
330+
looked too low and was right.
331+
332+
The second is the more useful lesson. The first review checked whether each claim was
333+
*true* without checking whether the *test measured the product* — the same failure
334+
`evaluator.py` had four days earlier. An adversarial review has to attack the
335+
measurement as well as the claim.
336+
337+
**And the plan contradicted the spec.** FR-007 promised English questions "no
338+
additional prompt content" while the plan always passes the language, which adds a
339+
sentence to every English prompt. FR-007 was the wrong half and was corrected to
340+
promise what is actually true: no extra model call, byte-identical retrieval query,
341+
and one added sentence stated rather than hidden.
342+
343+
### Worth watching
344+
345+
The French answer carried 2 `R-HSA` citations against the English answer's 9. One
346+
sample, so not a finding — but if non-English answers systematically cite less, this
347+
feature would be creating a quality gap while closing a language one. It is the kind
348+
of thing the ragas run in [spec 002](../002-default-llm-choice/spec.md) should look
349+
at once that harness is worth running.
350+
351+
### Not resolved
352+
353+
#125 also carried a hallucination grader and web-search wiring for Cross-Database.
354+
Neither is touched here, and closing that PR is not a decision about them; they
355+
belong to #123.

‎specs/004-answer-in-user-language/tasks.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -34,9 +34,9 @@ unasserted claim of that kind is worth nothing.
3434

3535
- [x] T010 Perturbation check: revert the call site and confirm the French test fails
3636
- [x] T011 Run `ruff check`, `ruff format --check`, `mypy`, `pytest`
37-
- [ ] T012 Close #140 with credit to @bleedblack1 — the target and the nomenclature rule are kept; the mechanism is not, with the measured number
38-
- [ ] T013 Close #125 with credit to @bhavyakeerthi3 for the mechanism, noting its hallucination-grading work belongs to #123 and is untouched
39-
- [ ] T014 Record the outcome in `specs/004-answer-in-user-language/spec.md`
37+
- [x] T012 Close #140 with credit to @bleedblack1 — the target and the nomenclature rule are kept; the mechanism is not, with the measured number
38+
- [x] T013 Close #125 with credit to @bhavyakeerthi3 for the mechanism, noting its hallucination-grading work belongs to #123 and is untouched
39+
- [x] T014 Record the outcome in `specs/004-answer-in-user-language/spec.md`
4040

4141
## Dependencies
4242

0 commit comments

Comments
 (0)