Skip to content

Commit 2eac8ed

Browse files
Continue in chat from a search-page answer
Spec 013, Story 2. The search page's AI answer can now be continued in a chat tab, the way analysis summaries can since #294. **Each answer gets its own ID, emitted in `done`.** It cannot be keyed by question: two readers of one search get different answers (~0.33 similarity run to run), and a question-keyed cache would let one continue the other's. `/api/answer` keeps each answered stream under a fresh `answer_id`, present in `done` only when `state` is `answered`, for an hour. What is kept is exactly what the page was sent -- the text after anchor and sources stripping, and the citations -- not the raw model output. `POST /api/handoff` now takes a discriminated request: `{"kind": "analysis", "token", "disclosure"}` or `{"kind": "search", "answer_id"}`. A search request carrying analysis fields is rejected rather than read as one. **No human-presence claim for a search handoff.** `/api/answer`, which produced the answer, deliberately does not require it -- public pathway text -- so the search page may have none to send. The analysis handoff still requires it, because it releases a reader's own analysis. Pinned in both directions: requiring presence for search fails one test, dropping it for analysis fails two. The handoff record is two types, `AnalysisHandoff` and `SearchHandoff`, rather than one with optional fields, so a search handoff cannot carry a disclosure tier that means nothing for it. mypy then found every place that had assumed a handoff was an analysis. The chat opens on the reader's own question as the human turn and the answer, with its cited sources, as the model's. Verified end to end through the real Turnstile gate as a first-time visitor, over HTTPS: a real answer from the real endpoint carried an answer_id; a handoff was minted with no human claim; the tab opened on the question and the same answer; "which protein kinase were we just discussing?" was answered "CDK5". The control -- the same question with no handoff -- could not say. Sabotage: storing anything other than the streamed text fails `test_what_is_kept_is_exactly_what_the_page_was_sent`. Two of the presence sabotages first came back void -- ruff had reformatted the conditional, so the anchor matched nothing -- and one printed "applied" over an unchanged file; both were redone with the change asserted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1 parent 04fc9c2 commit 2eac8ed

11 files changed

Lines changed: 491 additions & 50 deletions

File tree

‎bin/chat-chainlit.py‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -24,7 +24,7 @@
2424
run_analysis,
2525
)
2626
from handoff import seed
27-
from handoff.store import handoffs
27+
from handoff.store import AnalysisHandoff, handoffs
2828
from handoff.window import acknowledgement, claimed_id
2929
from util.chainlit_helpers import (
3030
PrefixedS3StorageClient,
@@ -229,8 +229,10 @@ async def continue_from_handoff(handoff_id: str) -> None:
229229
# `on_chat_start` sets `thread_id` from the session id. A claim arriving
230230
# before it has run would otherwise seed a thread called "None".
231231
thread_id: str = cl.user_session.get("thread_id") or cl.user_session.get("id")
232+
data = None
232233
try:
233-
data = await seed.analysis_data(handoff)
234+
if isinstance(handoff, AnalysisHandoff):
235+
data = await seed.analysis_data(handoff)
234236
seeded = await get_graph().seed_history(
235237
profile, thread_id=thread_id, messages=seed.seeded_turn(handoff, data)
236238
)
@@ -246,7 +248,11 @@ async def continue_from_handoff(handoff_id: str) -> None:
246248

247249
logger.info(
248250
"handoff claimed",
249-
extra={"tier": handoff.tier, "with_data": data is not None},
251+
extra={
252+
"kind": handoff.kind,
253+
"tier": getattr(handoff, "tier", None),
254+
"with_data": data is not None,
255+
},
250256
)
251257
await cl.Message(content=seed.shown_to_reader(handoff)).send()
252258

‎specs/010-search-page-answers/contracts/answer_endpoint.md‎

Lines changed: 11 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -59,11 +59,21 @@ event: citation
5959
data: {"st_id": "R-HSA-8862803", "display_name": "Deregulated CDK5 triggers..."}
6060
6161
event: done
62-
data: {"state": "answered", "seconds": 8.4}
62+
data: {"state": "answered", "seconds": 8.4, "answer_id": "Ev3z3JDmIIUF4gkKTfrn9VF83H"}
6363
```
6464

6565
`state` is one of `answered`, `nothing_found`, `refused`, `failed`.
6666

67+
**`answer_id` is present only when `state` is `answered`** (added 2026-09-25,
68+
spec 013). It names the answer exactly as this stream sent it -- the text after
69+
anchors and the sources list were stripped, and the citations -- kept for an
70+
hour so "Continue in chat" can open the chat on *that* answer rather than a
71+
regenerated one (the same question scores ~0.33 similarity run to run). Pass it
72+
to `POST /api/handoff` as `{"kind": "search", "answer_id": ...}`. It is keyed per
73+
answer, not per question, so two readers of the same search never share one.
74+
Every other `done` has exactly `state` and `seconds`; a caller that ignores
75+
unknown keys parses them all the same way.
76+
6777
> **This endpoint does not require a human-presence claim and
6878
> `/api/analysis-summary` does.** A caller token without one gets a full
6979
> answer here and `no_human` there. Deliberate: this returns public pathway

‎specs/013-continue-chat-hand/spec.md‎

Lines changed: 20 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44

55
**Created**: 2026-09-25
66

7-
**Status**: Story 1 built and verified end to end; Stories 2 and 3 not started
7+
**Status**: Stories 1 and 2 built and verified end to end; Story 3 needs a shared store
88

99
**Input**: Adam, 2026-09-25: "with the chat on both the search page and the analysis results we want there to be a button to go to the [chat] interface … two options … either the logged in version or the guest version of the chat app and have the context already be set with either the search results or the analysis summary with the summary data." And: "the chat should open in a new tab."
1010

@@ -165,3 +165,22 @@ on "couldn't load the summary". Two ways to do Story 3, to decide later:
165165
- the website calls the logged-in deployment's own summary and handoff
166166
endpoints -- simpler, but the summary would be generated a second time in
167167
that process, and would not be the one the reader saw (FR-002).
168+
169+
## Story 2, 2026-09-25
170+
171+
**Search-page answers can be continued.** `/api/answer` now keeps each answered
172+
stream under an `answer_id`, emitted in `done` only when `state` is
173+
`answered`, and `POST /api/handoff` accepts `{"kind": "search", "answer_id"}`.
174+
Keyed per answer, never per question: two readers of one search get different
175+
answers, and neither may continue the other's. What is kept is what the page
176+
was sent -- after anchor and sources stripping -- not the raw model output.
177+
178+
**No human-presence claim for a search handoff**, unlike an analysis one.
179+
`/api/answer` does not require it either (public pathway text), so the search
180+
page may have none to send. Pinned in both directions: requiring it for search
181+
fails one test, dropping it for analysis fails two.
182+
183+
Verified end to end through the real Turnstile gate as a first-time visitor:
184+
a real answer, a handoff minted with no human claim, the tab opens on the
185+
question and the same answer, and "which protein kinase were we just
186+
discussing?" is answered "CDK5". The control, with no handoff, cannot say.

‎src/api/answer.py‎

Lines changed: 30 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -29,6 +29,7 @@
2929
from pydantic import BaseModel, Field
3030

3131
from agent.registry import get_graph
32+
from api.answer_store import StoredAnswer, answers
3233
from util.anchor_strip import AnchorStripper
3334
from util.caller_token import TokenRejectedError, verify
3435
from util.logging import logging
@@ -129,6 +130,11 @@ async def stream() -> AsyncIterator[str]:
129130
# worse one, since `AnchorStripper` has just taken its links off.
130131
sources = SourcesSectionStripper()
131132
tokens_sent = 0
133+
# What the reader is shown, kept so "Continue in chat" can open the
134+
# chat on this answer rather than a regenerated one (spec 013). The
135+
# text after stripping, because that is what the page renders.
136+
shown: list[str] = []
137+
cited: list[tuple[str, str]] = []
132138
try:
133139
async with asyncio.timeout(ANSWER_TIMEOUT_SECONDS):
134140
async for event in graph.astream_answer(
@@ -147,6 +153,7 @@ async def stream() -> AsyncIterator[str]:
147153
text = sources.feed(stripper.feed(event.text))
148154
if text:
149155
tokens_sent += 1
156+
shown.append(text)
150157
yield _sse("token", {"text": text})
151158
elif event.kind == "citation":
152159
# Exactly one identifier, never both and never an empty
@@ -159,6 +166,9 @@ async def stream() -> AsyncIterator[str]:
159166
if event.st_id
160167
else {"url": event.url}
161168
)
169+
cited.append(
170+
(event.st_id or event.url or "", event.display_name or "")
171+
)
162172
yield _sse(
163173
"citation",
164174
{**identifier, "display_name": event.display_name},
@@ -209,9 +219,26 @@ async def stream() -> AsyncIterator[str]:
209219
state = "failed"
210220
held = sources.feed(stripper.flush()) + sources.flush()
211221
if held:
222+
shown.append(held)
212223
yield _sse("token", {"text": held})
213-
yield _sse(
214-
"done", {"state": state, "seconds": round(time.monotonic() - started, 1)}
215-
)
224+
done: dict[str, Any] = {
225+
"state": state,
226+
"seconds": round(time.monotonic() - started, 1),
227+
}
228+
if state == "answered":
229+
# Only an answer can be continued. The ID, not the question, is
230+
# the key: two readers asking the same thing get different text.
231+
answer_id = answers.put(
232+
StoredAnswer(
233+
question=body.question,
234+
text="".join(shown),
235+
citations=tuple(cited),
236+
release=release,
237+
created_at=time.time(),
238+
)
239+
)
240+
if answer_id:
241+
done["answer_id"] = answer_id
242+
yield _sse("done", done)
216243

217244
return StreamingResponse(stream(), media_type="text/event-stream")

‎src/api/answer_store.py‎

Lines changed: 72 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,72 @@
1+
"""Answers the search page was shown, kept briefly so they can be continued.
2+
3+
Spec 013, Story 2. "Continue in chat" beside a search-page answer must open
4+
the chat on *that* answer. It cannot be regenerated: the same question
5+
through the same endpoint scores ~0.33 similarity run to run, so the reader
6+
would be greeted with different text from the one they clicked from.
7+
8+
**Keyed by a per-answer ID, not by the question.** Two readers asking the
9+
same question get different answers; a cache keyed by question would let
10+
one reader continue another's. The endpoint issues an ID in its `done`
11+
event, and that is what the website hands back.
12+
13+
**What is kept is exactly what the reader saw** -- the text after anchors and
14+
the trailing sources list were stripped, and the citations as sent -- not
15+
the raw model output.
16+
17+
Bounded and process-local, like the analysis summary cache. An answer is
18+
public pathway text, so keeping it briefly carries no disclosure; the window
19+
only has to outlast a reader deciding to click.
20+
"""
21+
22+
import secrets
23+
import time
24+
from collections import OrderedDict
25+
from dataclasses import dataclass, field
26+
27+
DEFAULT_TTL_SECONDS = 60 * 60
28+
DEFAULT_MAX_ENTRIES = 2048
29+
30+
31+
@dataclass(frozen=True)
32+
class StoredAnswer:
33+
question: str
34+
text: str
35+
#: `(identifier, display_name)`, where identifier is an st_id or a URL.
36+
citations: tuple[tuple[str, str], ...]
37+
release: int | None
38+
created_at: float
39+
40+
41+
@dataclass
42+
class AnswerStore:
43+
ttl_seconds: float = DEFAULT_TTL_SECONDS
44+
max_entries: int = DEFAULT_MAX_ENTRIES
45+
_entries: OrderedDict[str, StoredAnswer] = field(default_factory=OrderedDict)
46+
47+
def put(self, answer: StoredAnswer) -> str | None:
48+
"""Keep an answer; returns its ID, or None if there was nothing to keep.
49+
50+
An empty answer is never stored: continuing it would open a chat on
51+
nothing, and say it was the reader's answer.
52+
"""
53+
if not answer.text.strip():
54+
return None
55+
answer_id = secrets.token_urlsafe(18)
56+
self._entries[answer_id] = answer
57+
while len(self._entries) > self.max_entries:
58+
self._entries.popitem(last=False)
59+
return answer_id
60+
61+
def get(self, answer_id: str, *, now: float | None = None) -> StoredAnswer | None:
62+
found = self._entries.get(answer_id)
63+
if found is None:
64+
return None
65+
if (time.time() if now is None else now) - found.created_at > self.ttl_seconds:
66+
del self._entries[answer_id]
67+
return None
68+
return found
69+
70+
71+
#: Written by `/api/answer`, read by `/api/handoff`.
72+
answers = AnswerStore()

‎src/api/handoff.py‎

Lines changed: 67 additions & 24 deletions
Original file line numberDiff line numberDiff line change
@@ -25,7 +25,7 @@
2525

2626
import os
2727
import time
28-
from typing import Any
28+
from typing import Annotated, Any, Literal
2929

3030
from fastapi import APIRouter, Request
3131
from fastapi.responses import JSONResponse
@@ -34,7 +34,14 @@
3434
from analysis.client import current_release
3535
from analysis.disclosure import Tier
3636
from api.analysis_summary import stored_summary
37-
from handoff.store import DEFAULT_TTL_SECONDS, Handoff, handoffs
37+
from api.answer_store import StoredAnswer, answers
38+
from handoff.store import (
39+
DEFAULT_TTL_SECONDS,
40+
AnalysisHandoff,
41+
Handoff,
42+
SearchHandoff,
43+
handoffs,
44+
)
3845
from util.caller_token import (
3946
TokenRejectedError,
4047
human_presence_detail,
@@ -51,19 +58,36 @@
5158
_limiter = limiter_from_env()
5259

5360

54-
class HandoffRequest(BaseModel):
55-
kind: str = Field(pattern="^analysis$")
61+
class AnalysisHandoffRequest(BaseModel):
62+
kind: Literal["analysis"]
5663
token: str = Field(min_length=1, max_length=256)
5764
#: The tier of the summary the reader was shown. Required, no default.
5865
disclosure: Tier
5966
caller_token: str = ""
6067

6168

69+
class SearchHandoffRequest(BaseModel):
70+
kind: Literal["search"]
71+
#: From the `done` event of the `/api/answer` stream the page rendered.
72+
answer_id: str = Field(min_length=1, max_length=128)
73+
caller_token: str = ""
74+
75+
76+
HandoffRequest = Annotated[
77+
AnalysisHandoffRequest | SearchHandoffRequest, Field(discriminator="kind")
78+
]
79+
80+
6281
def _refuse(status: int, reason: str, log: str) -> JSONResponse:
6382
logger.info("handoff refused", extra={"reason": reason, "detail": log})
6483
return JSONResponse(status_code=status, content={"reason": reason})
6584

6685

86+
def stored_answer(answer_id: str) -> StoredAnswer | None:
87+
"""An answer `/api/answer` showed a reader, if it is still kept."""
88+
return answers.get(answer_id)
89+
90+
6791
def own_chat_path() -> str:
6892
"""Where this process's chat is mounted, e.g. `/chat/guest`."""
6993
return (os.getenv("CHAINLIT_URI") or "/chat").rstrip("/")
@@ -79,7 +103,13 @@ async def create_handoff(body: HandoffRequest, request: Request) -> JSONResponse
79103
except TokenRejectedError as rejected:
80104
return _refuse(403, "no_caller", rejected.reason)
81105

82-
presence = human_presence_reason(claims, time.time())
106+
# Human presence for an analysis only. It releases a reader's own
107+
# analysis summary; a search answer is public pathway text, and
108+
# `/api/answer` -- which produced it -- deliberately does not require
109+
# presence either, so the search page may not have the claim to send.
110+
presence = (
111+
human_presence_reason(claims, time.time()) if body.kind == "analysis" else None
112+
)
83113
if presence:
84114
detail = (
85115
human_presence_detail(claims, time.time())
@@ -93,25 +123,37 @@ async def create_handoff(body: HandoffRequest, request: Request) -> JSONResponse
93123
if not _limiter.allow(key or identity_of(claims, body.caller_token)):
94124
return _refuse(429, "rate_limited", "rate limited")
95125

96-
release = await current_release()
97-
if not release:
98-
# Summaries are cached per release, so without knowing the release
99-
# there is no way to find the one the reader saw. Refusing is the
100-
# honest answer; guessing a release could hand off a summary of a
101-
# result the Analysis Service has since deleted.
102-
return _refuse(503, "no_release", "current release unknown")
103-
stored = stored_summary(body.token, release, body.disclosure)
104-
if stored is None:
105-
# Not generated, evicted, from a previous release, or requested at a
106-
# tier the reader never chose. In every case there is nothing the
107-
# reader has seen to continue, and generating one here would break
108-
# both FR-002 and FR-003.
109-
return _refuse(
110-
404, "no_summary", f"no stored summary at tier {body.disclosure}"
126+
handoff: Handoff
127+
if isinstance(body, SearchHandoffRequest):
128+
answer = stored_answer(body.answer_id)
129+
if answer is None:
130+
# Unknown, expired, or never answered: nothing the reader saw.
131+
return _refuse(404, "no_answer", "no stored answer for that id")
132+
handoff = SearchHandoff(
133+
kind="search",
134+
question=answer.question,
135+
summary=answer.text,
136+
citations=answer.citations,
137+
created_at=time.time(),
111138
)
112-
113-
handoff_id = handoffs.put(
114-
Handoff(
139+
else:
140+
release = await current_release()
141+
if not release:
142+
# Summaries are cached per release, so without knowing the release
143+
# there is no way to find the one the reader saw. Refusing is the
144+
# honest answer; guessing a release could hand off a summary of a
145+
# result the Analysis Service has since deleted.
146+
return _refuse(503, "no_release", "current release unknown")
147+
stored = stored_summary(body.token, release, body.disclosure)
148+
if stored is None:
149+
# Not generated, evicted, from a previous release, or requested at
150+
# a tier the reader never chose. In every case there is nothing the
151+
# reader has seen to continue, and generating one here would break
152+
# both FR-002 and FR-003.
153+
return _refuse(
154+
404, "no_summary", f"no stored summary at tier {body.disclosure}"
155+
)
156+
handoff = AnalysisHandoff(
115157
kind="analysis",
116158
token=body.token,
117159
release=release,
@@ -120,7 +162,8 @@ async def create_handoff(body: HandoffRequest, request: Request) -> JSONResponse
120162
citations=stored.citations,
121163
created_at=time.time(),
122164
)
123-
)
165+
166+
handoff_id = handoffs.put(handoff)
124167
payload: dict[str, Any] = {
125168
"id": handoff_id,
126169
"expires_in": int(DEFAULT_TTL_SECONDS),

0 commit comments

Comments
 (0)