Strip the trailing source list, and stop the website guessing our heading - #258
Merged
Merged
Conversation
…instead The website strips a trailing source list from our prose before rendering, because it shows citations as its own chips and a markdown list underneath reads as a duplicate. It matched `## Sources`, then `## Most relevant sources`, then `relevant references`, `Key sources` and `Top citations`, and asked which heading level the contract specifies so it could stop guessing. It specifies none. All three answer prompts say only "provide a bullet-point list of each unique citation anchor" and say nothing about a heading, so the model invents one per answer. Those five phrasings are not a contract, they are five samples from an unconstrained generator, and no pattern written against them can win. Both halves are ours to fix, so both are here. The prompts now pin the heading to exactly `## Sources`. And the answer endpoint drops that section before it reaches the caller, so no client has to match anything: the `citation` events are the authoritative list, and the prose copy is strictly worse -- `AnchorStripper` has just taken its links off, leaving bare display names. `SourcesSectionStripper` still accepts the observed variants, because a prompt is an instruction and not a guarantee, but it is bounded to a heading or bold-only line of at most five words naming sources, rather than any line mentioning one. The tests use the real headings the website saw. Two things went wrong while writing it, both worth the tests they left behind. Holding back every unterminated line to see whether it became a heading turned the answer into a single blob at flush, because prose often has no newline until it ends. Two existing endpoint tests caught it. Only a partial line that could still open a heading is held now. And `^` under re.MULTILINE matches at offset zero whether or not that offset is a line start, so once text had been released, a bolded word mid-sentence looked like a heading and ate the rest of the answer. The stripper tracks whether its buffer begins at a real line start. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two defects, both found by attacking the change rather than re-reading it, and both in the direction that costs the reader the rest of the answer. **A heading about biology matched.** The pattern accepted any heading of five words or fewer that mentioned sources, so "## Sources of reactive oxygen species" -- an entirely plausible Reactome heading -- truncated the answer there. Four of seven realistic headings did. The bound was on length, not on meaning, which is not a bound on anything that matters. The noun must now be the last word of the heading, and only a closed set of qualifiers may precede it. "Most relevant sources" is a source list; "Sources of oxidative stress" is biology; "Cellular sources" is left alone because "cellular" is not a qualifier we have ever seen on a source list. **A prefix of a longer heading matched mid-stream.** Fed character by character, "## Sources" matches before " of reactive oxygen species" arrives, so the stripper committed and dropped the rest. The whole-string probe passed and the character-by-character one failed, which is the only reason this was caught: a heading is not decided until its line is terminated, or the stream ends. The asymmetry is now stated where the pattern is defined, because it is the thing that should settle any future widening. A heading this misses costs a duplicate list at the end -- cosmetic, and what the website lives with today. A heading this matches wrongly costs the answer. When in doubt, do not match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
adamjohnwright
force-pushed
the
fix/strip-trailing-source-list
branch
from
September 19, 2026 03:45
5d34113 to
110e0a3
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The website session asked which heading level our contract specifies for the source list at the end of an answer, so it could stop widening its strip pattern. It had already matched
## Sources,## Most relevant sources,relevant references,Key sourcesandTop citations.The contract specifies none. All three answer prompts say only:
No heading level, no heading text, no instruction to emit a heading at all. Those five phrasings are not five readings of a contract — they are five samples from an unconstrained generator, and no pattern written against them can win. The fix is ours, in both halves.
Pin the heading
The prompts now require exactly
## Sources, and say why, so the instruction survives someone editing the prompt later.Strip the section server-side
SourcesSectionStripperdrops the heading and everything after it on the answer endpoint. No client needs to match anything: thecitationevents are the authoritative list, and the prose copy is strictly worse —AnchorStripperhas just removed its links, leaving bare display names.It still accepts the observed variants, because a prompt is an instruction and not a guarantee, but it is bounded to a heading or bold-only line of at most five words naming sources — not any line mentioning one. The tests use the real headings the website observed, and two tests pin the direction that would actually hurt: prose about sources, and a long bold sentence ending in "sources", must survive untouched.
Two bugs found while writing it
Holding every unterminated line destroyed streaming. Prose often has no newline until it ends, so the whole answer arrived as one blob at
flush. Two existing endpoint tests caught it — the stripper now holds back only a partial line that could still open a heading, and there is a direct test for it.^underre.MULTILINEmatches at offset zero whether or not that offset is a line start. Once text had been released the buffer began mid-sentence, so a bolded word in prose looked like a heading and ate the rest of the answer. The stripper tracks whether its buffer begins at a real line start.Checks
483 passed, 1 skipped. ruff, ruff format and mypy clean.
Contract updated in
specs/010-search-page-answers/contracts/answer_endpoint.md, including an explicit instruction that callers must not pattern-match the heading. The website session has agreed to delete its patterns once this is on beta rather than keep them as a safety net.🤖 Generated with Claude Code