Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions specs/010-search-page-answers/contracts/answer_endpoint.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,28 @@ linked phrase in the middle. Where the model used a link as a trailing citation,
that leaves the pathway title as a bare clause -- cosmetic, and the structured
`citation` events are the reliable source for links.

**No trailing source list.** The answer prompts ask for a bullet list of every
citation at the end, which the chat UI renders. This caller gets its links from
the `citation` events, so that list is a duplicate -- and after the anchors come
off, a duplicate with no links in it. The endpoint drops the heading and
everything after it.

Callers **must not** pattern-match the heading themselves. Until 2026-09-19 the
prompts specified no heading at all, only "a bullet-point list of each unique
citation anchor", so the model invented one per answer: `## Sources`,
`## Most relevant sources`, `relevant references`, `Key sources`,
`Top citations`. The website was matching those and could not win, because it
was fitting samples from an unconstrained generator. The prompts now pin the
heading to exactly:

```
## Sources
```

and `SourcesSectionStripper` removes it on the served path. The stripper still
accepts the older variants, because a prompt is an instruction and not a
guarantee, but no caller needs to know that.

## Properties worth holding to

**It must be safe to ignore.** Any failure, timeout, refusal or unverified caller
Expand Down
9 changes: 7 additions & 2 deletions src/api/answer.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@
from util.caller_token import TokenRejectedError, verify
from util.logging import logging
from util.rate_limit import identity_of, limiter_from_env
from util.sources_section import SourcesSectionStripper

logger = logging.getLogger(__name__)

Expand Down Expand Up @@ -123,6 +124,10 @@ async def stream() -> AsyncIterator[str]:
# separate events, so they come out here -- across fragment boundaries,
# because one anchor arrives as twenty-odd fragments.
stripper = AnchorStripper()
# And the trailing source list goes too: this caller renders citations
# from the `citation` events, so the prose copy is a duplicate -- and a
# worse one, since `AnchorStripper` has just taken its links off.
sources = SourcesSectionStripper()
tokens_sent = 0
try:
async with asyncio.timeout(ANSWER_TIMEOUT_SECONDS):
Expand All @@ -139,7 +144,7 @@ async def stream() -> AsyncIterator[str]:
enable_postprocess=False,
):
if event.kind == "token":
text = stripper.feed(event.text)
text = sources.feed(stripper.feed(event.text))
if text:
tokens_sent += 1
yield _sse("token", {"text": text})
Expand Down Expand Up @@ -202,7 +207,7 @@ async def stream() -> AsyncIterator[str]:
# terminal event to stop waiting.
logger.exception("answering %r failed", body.question[:80])
state = "failed"
held = stripper.flush()
held = sources.feed(stripper.flush()) + sources.flush()
if held:
yield _sse("token", {"text": held})
yield _sse(
Expand Down
4 changes: 4 additions & 0 deletions src/retrievers/plantreactome/prompt.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,10 @@
- Use accessible language while maintaining technical precision.
- Ensure the narrative flows logically, presenting background, mechanisms, and significance
5. Source list at the end: After the main narrative, provide a bullet-point list of each unique citation anchor exactly once, in the same <a href="URL">Node Name</a> format.
- Head the list with exactly this line and nothing else: `## Sources`
Not a variation on it. The search page strips this section by that
exact heading, because it renders the citations itself; a different
wording leaves the reader a duplicate list.
- Examples:
- <a href="https://plantreactome.gramene.org/content/detail/R-OSA-9640713">Mitosis</a>
- <a href="https://plantreactome.gramene.org/content/detail/R-OSA-9640670">Cell Cycle</a>
Expand Down
4 changes: 4 additions & 0 deletions src/retrievers/reactome/prompt.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,10 @@
- Use accessible language while maintaining technical precision.
- Ensure the narrative flows logically, presenting background, mechanisms, and significance
5. Source list at the end: After the main narrative, provide a bullet-point list of each unique citation anchor exactly once, in the same <a href="URL">Node Name</a> format.
- Head the list with exactly this line and nothing else: `## Sources`
Not a variation on it. The search page strips this section by that
exact heading, because it renders the citations itself; a different
wording leaves the reader a duplicate list.
- Examples:
- <a href="https://reactome.org/content/detail/R-HSA-109581">Apoptosis</a>
- <a href="https://reactome.org/content/detail/R-HSA-1640170">Cell Cycle</a>
Expand Down
4 changes: 4 additions & 0 deletions src/retrievers/userguide/prompt.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,10 @@
- Use accessible language; avoid unnecessary jargon.
- Prefer numbered steps for multi-step procedures.
5. Source list at the end: After the main answer, provide a bullet-point list of each unique citation anchor exactly once, in the same <a href="URL">display_name</a> format.
- Head the list with exactly this line and nothing else: `## Sources`
Not a variation on it. The search page strips this section by that
exact heading, because it renders the citations itself; a different
wording leaves the reader a duplicate list.
- Examples:
- <a href="https://reactome.org/userguide/pathway-browser">Pathway Browser</a>
- <a href="https://reactome.org/userguide/searching">Searching Reactome</a>
Expand Down
148 changes: 148 additions & 0 deletions src/util/sources_section.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,148 @@
"""Drop the trailing source list from a token stream split at arbitrary points.

The search page renders citations as its own chips, built from the `citation`
events, so the prose list at the end of an answer is a strictly worse copy of
data the caller already has -- worse still after `AnchorStripper`, which leaves
it as bare display names with no links.

The website was matching the heading itself and could not win. Measured
2026-09-19, our prompts never specified one: they asked for "a bullet-point
list of each unique citation anchor" and said nothing about a heading, so the
model invented one per answer -- `## Sources`, `## Most relevant sources`,
`relevant references`, `Key sources`, `Top citations`. That is not five
phrasings of a contract, it is five samples from an unconstrained generator.

The prompts now pin the heading to `## Sources`. This is the second half: strip
it here so no caller has to pattern-match model output at all. The pattern below
is still tolerant, because the prompt is an instruction and not a guarantee, but
it is bounded -- a heading or bold-only line, at most five words, naming sources
-- rather than any line mentioning the word.
"""

import re

# A whole line that is a markdown heading or a bold-only line naming nothing
# but the source list. Two rules keep it off real content, and the first was
# learned the hard way: an earlier version allowed any short heading that
# mentioned sources, and "## Sources of reactive oxygen species" -- an entirely
# plausible Reactome heading -- truncated the answer there.
#
# 1. The noun must be the LAST word. "Sources of oxidative stress" is about
# biology; "Most relevant sources" is a source list.
# 2. Only a closed set of qualifiers may precede it, so "Cellular sources" is
# left alone.
#
# The asymmetry justifies the strictness. A heading this misses costs the
# reader a duplicate list at the end -- cosmetic, and exactly what the website
# lives with today. A heading this matches wrongly costs them the rest of the
# answer. When in doubt, do not match.
_QUALIFIER = (
r"(?:most|more|relevant|key|top|main|primary|all|further|additional"
r"|related|supporting|complete|full|cited|used)"
)
_HEADING = re.compile(
r"^[ \t]*(?:#{1,6}[ \t]+|\*\*[ \t]*)"
rf"(?:{_QUALIFIER}[ \t]+){{0,3}}"
r"(?:sources?|references?|citations?)"
r"[ \t]*:?[ \t]*(?:\*\*)?[ \t]*:?[ \t]*$",
re.IGNORECASE | re.MULTILINE,
)

# Held back while a partial line could still turn out to be that heading. A
# heading is one short line; beyond this the text is prose and is released.
_MAX_HELD = 120


def _could_become_heading(partial: str) -> bool:
"""True while an unterminated line might still turn into the heading.

A heading starts a line with `#` or `**`. A bullet (`* `) cannot, and
neither can prose, so both stream straight through.
"""
stripped = partial.lstrip(" \t")
if len(stripped) > _MAX_HELD:
return False # Too long to be a heading; it is prose.
if not stripped:
return True
if stripped[0] == "#":
return True
return stripped == "*" or stripped.startswith("**")


class SourcesSectionStripper:
"""Feed fragments in, get the answer without its trailing source list.

Everything from the heading to the end of the stream is dropped: the
prompts place the list last, so there is nothing after it to keep.
"""

def __init__(self) -> None:
self._buffer = ""
self._done = False
# Whether the buffer currently begins at a real start of line. Once
# text has been released, position 0 is mid-line -- and `^` under
# re.MULTILINE matches there anyway, which turned a bolded word in
# mid-sentence into a heading and ate the rest of the answer.
self._at_line_start = True

def feed(self, text: str) -> str:
if self._done:
return ""
self._buffer += text
# Terminated only: mid-stream, "## Sources" matches before
# " of reactive oxygen species" has arrived, and deciding then drops
# the rest of a perfectly good answer. Only the whole-string probe
# caught this; the character-by-character one is what found it.
match = self._find_heading(terminated_only=True)
if match:
out = self._buffer[: match.start()]
self._buffer = ""
self._done = True
return out
# Hold back only a partial line that could still become that heading.
# Holding every unterminated line instead would defeat the point of
# the endpoint: an answer often has no newline until it ends, so the
# whole thing would arrive in one blob at `flush`. Two existing tests
# caught exactly that.
newline = self._buffer.rfind("\n")
if newline == -1 and not self._at_line_start:
cut = len(self._buffer) # No line start in here at all.
else:
cut = newline + 1
if not _could_become_heading(self._buffer[cut:]):
cut = len(self._buffer)
out, self._buffer = self._buffer[:cut], self._buffer[cut:]
if out:
self._at_line_start = out.endswith("\n")
return out

def flush(self) -> str:
"""Whatever is still held, once the stream has ended."""
if self._done:
return ""
match = self._find_heading(terminated_only=False)
held = self._buffer[: match.start()] if match else self._buffer
self._buffer = ""
self._done = True
return held

def _find_heading(self, *, terminated_only: bool) -> re.Match[str] | None:
"""The heading, ignoring a match at offset 0 when that is mid-line.

`terminated_only` rejects a match that runs to the end of the buffer,
because more of that line may still arrive. At `flush` the stream has
ended, so there is nothing more to wait for.
"""
position = 0
while True:
match = _HEADING.search(self._buffer, position)
if match is None:
return None
if terminated_only and match.end() >= len(self._buffer):
return None # The line has not ended; it may yet grow.
if match.start() or self._at_line_start:
return match
newline = self._buffer.find("\n")
if newline == -1:
return None
position = newline + 1
41 changes: 41 additions & 0 deletions tests/api/test_answer_endpoint.py
Original file line number Diff line number Diff line change
Expand Up @@ -488,3 +488,44 @@ async def drive() -> None:
assert any(
"abandoned by the caller" in message for message in messages
), f"no record of the abandoned stream; logged: {messages}"


def test_the_trailing_source_list_never_reaches_the_caller(
keys: tuple[str, str], monkeypatch: pytest.MonkeyPatch
) -> None:
# The website renders citations as chips from the `citation` events, so the
# prose list at the end is a duplicate -- and after AnchorStripper has taken
# the links off, a worse one. They were matching the heading and could not
# win: our prompts never specified one, so the model invented a different
# heading per answer. Stripped here instead, on the served path.
stub = _StubGraph(
[
AnswerEvent(kind="citation", st_id="R-HSA-1", display_name="Apoptosis"),
AnswerEvent(kind="token", text="CDK5 phosphorylates tau.\n\n"),
AnswerEvent(kind="token", text="## Most rele"),
AnswerEvent(kind="token", text="vant sources\n- "),
AnswerEvent(
kind="token",
text='<a href="https://reactome.org/content/detail/R-HSA-1">Apoptosis</a>\n',
),
AnswerEvent(kind="done", state="answered"),
]
)
monkeypatch.setattr("api.answer.get_graph", lambda *_a, **_k: stub)
private_pem, public_pem = keys
response = _client(public_pem).post(
f"{PREFIX}/answer",
json={"question": "what does CDK5 do?", "caller_token": _token(private_pem)},
)
assert response.status_code == 200
prose = "".join(
json.loads(data)["text"]
for kind, data in _events(response.text)
if kind == "token"
)
assert prose.strip() == "CDK5 phosphorylates tau."
assert "sources" not in prose.lower()
assert "Apoptosis" not in prose, "the citation event is the list, not the prose"
# The citation itself must survive: stripping the prose copy must not cost
# the caller the data it renders chips from.
assert any(kind == "citation" for kind, _ in _events(response.text))
Loading
Loading