Retry on what happened, not on what the answer said about it - #270
Merged
Merged
Conversation
T023. The sweep decided whether to retry a question by matching the answer's prose, and one of the markers was "could not find out" -- the wording `src/reactome_mcp/answer.py` *instructs* the model to use for a legitimate empty result. So the marker matched correct answers by construction, and since a match triggers a retry, a real regression got a second attempt and could pass. The gate forgave precisely what it exists to catch. Matching further down the pipe could never fix that. The tool exception is caught, logged, stringified into the model's context and then paraphrased; by the time anything reads prose, "a lookup failed" and "there is genuinely nothing" are the same sentence. The fix belongs where the information is discarded. `answer_from_live_services` now takes an optional `LiveReport` and records the exception before it becomes text. The graph carries it out as `live_tool_failed`, and the sweep retries on that. `TRANSIENT` and `_looks_transient` are deleted, with a test asserting they do not come back: if prose matching returns, it will return as a list of markers, and that is cheap to detect. The report is optional, so none of the twelve existing call sites changed. Both halves verified by sabotage. Making the retry ignore the signal fails the retry test; making the report never populate fails the recording test. And the retried case and the not-retried case now use the *same* answer text, so only the signal separates them -- which is the property that was missing before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
T023, and it closes the loop on something the website session and I both flagged: an exception that is caught, logged, stringified into a prompt and paraphrased has been through a lossy channel by design. Matching on the far end recovers nothing.
The bug being fixed
The sweep decided whether to retry by matching the answer's prose, and one marker was
"could not find out"— which is the wordingsrc/reactome_mcp/answer.pyinstructs the model to use for a legitimate empty result:So the marker matched correct answers by construction. Since a match triggers a retry, a real regression got a second attempt and could pass. The gate forgave exactly what it exists to catch.
The fix, where the information is discarded
answer_from_live_servicestakes an optionalLiveReportand records the tool exception before stringifying it into the model's context — the last point at which it is still a fact rather than a paraphrase. The graph carries it out aslive_tool_failed; the sweep retries on that.TRANSIENTand_looks_transientare deleted, with a test asserting they do not come back. If prose matching returns it will return as a list of markers, and that is cheap to detect.The report is optional, so none of the twelve existing call sites changed.
Verified by sabotage
And the retried case and the not-retried case now use the same answer text —
"I could not find out."— so only the signal separates them. That property is what was missing before, and it is why the old test passed against the broken behaviour.517 passed, 1 skipped; mypy over 139 files, ruff and ruff format clean.
🤖 Generated with Claude Code