From 09629df18da92a2eaac221a334ec55da48e3c73c Mon Sep 17 00:00:00 2001 From: Adam Wright Date: Thu, 17 Sep 2026 15:39:41 +0000 Subject: [PATCH 1/2] Spec the search-page answers, and the number that decides the design Adam wants chatbot answers in the website search results, Google-overview style, for verified humans only, coordinated with the website repo. The thing worth knowing before anyone designs anything: a whole answer takes 25.4s and 32.2s on two runs of the same question, measured today on beta against Release97. Retrieval is 14.9s of it. A Google AI overview arrives in one to two seconds. At twenty-five a panel spins for the entire time someone reads the ordinary results and leaves. So the spec states the budget as a requirement -- first token in 2s, complete in 10s -- rather than leaving it to be discovered in implementation. Two things follow: the answer has to stream, and spec 009 is on this feature's critical path rather than a parallel nicety, because retrieval is the largest single component. Three things this repo does not have: an answer endpoint at all (chat-fastapi.py serves the captcha pages and a landing page; Chainlit owns the conversation over websockets), a way for the website to present proof that it verified a person, and any rule about which searches deserve an answer. The last matters because search pages are crawled and every crawled search reaching the model is an unbounded bill. Turnstile already exists here, and the bug fixed this morning is directly relevant: a deployment mounting CLOUDFLARE_SECRET_KEY as a Docker secret had the captcha silently disabled. That is the control this feature depends on. contracts/answer_endpoint.md is the part the website session can start from. It commits to a wire shape and, more importantly, to properties: refuse before any model call, fail invisibly so the search page never breaks, citations as stable IDs so the site styles its own links, and one graph behind both the endpoint and the chat so they cannot disagree. Scope is this repo only. The panel, its design, and the wider "chat integrated into the site's feel" work are the website's, and deserve their own spec once this contract exists. Co-Authored-By: Claude Opus 5 --- .../contracts/answer_endpoint.md | 88 ++++++++ specs/010-search-page-answers/spec.md | 193 ++++++++++++++++++ 2 files changed, 281 insertions(+) create mode 100644 specs/010-search-page-answers/contracts/answer_endpoint.md create mode 100644 specs/010-search-page-answers/spec.md diff --git a/specs/010-search-page-answers/contracts/answer_endpoint.md b/specs/010-search-page-answers/contracts/answer_endpoint.md new file mode 100644 index 0000000..94778a7 --- /dev/null +++ b/specs/010-search-page-answers/contracts/answer_endpoint.md @@ -0,0 +1,88 @@ +# Contract: the search-page answer endpoint + +What this repo commits to providing, and what the website repo can build against. +Shapes are proposals until both sides agree; the *properties* below are the parts +worth arguing about. + +## Request + +``` +POST /chat/api/answer +Content-Type: application/json + +{ "question": "what does CDK5 phosphorylate in Alzheimer disease?", + "human_token": "" } +``` + +`human_token` is D1 and **not yet decided** -- a signed cookie on the shared parent +domain, a short-lived minted token, or a server-side vouch. This repo verifies +evidence; it does not perform the check. + +## Response: Server-Sent Events + +Streaming rather than a single JSON body, because the whole answer currently takes +25-32s and a search page cannot wait. Streaming turns that into "something appears in +about two seconds". + +``` +event: start +data: {"release": 97, "answered": true} + +event: token +data: {"text": "CDK5, when bound to p25, phosphorylates "} + +event: citation +data: {"st_id": "R-HSA-8862803", "display_name": "Deregulated CDK5 triggers..."} + +event: done +data: {"state": "answered", "seconds": 8.4} +``` + +`state` is one of `answered`, `nothing_found`, `refused`, `failed`. + +### Why citations are separate events + +So the website renders links in its own style. Returning prose with embedded HTML +anchors -- what the chat UI does today -- would force the search page to parse them +back out and re-style them. Stable IDs resolve at +`reactome.org/content/detail/`. + +## Properties worth holding to + +**It must be safe to ignore.** Any failure, timeout, refusal or unverified caller +produces a `done` with a non-`answered` state. The website renders no panel. The +search page must never be slower or broken because this service is down (FR-006, +SC-004). + +**No answer without a token.** Refused before any model call, not after (FR-003). +Search pages get crawled, and every crawled search reaching the model is a bill. + +**Not every search gets a panel.** The caller may ask for every query; this service +decides, and may answer `nothing_found` because the question is navigational rather +than answerable (D2, User Story 3). + +**The answer is tied to a release.** `start` carries it so a cached answer can be +invalidated after a release (FR-007). + +**One brain.** The endpoint and the chat UI share a graph. If they can disagree about +the same question, that is a defect. + +## Budget + +| | target | today | +|---|---|---| +| first token | 2s | n/a -- no streaming endpoint exists | +| complete | 10s | **25.4s and 32.2s** measured 2026-09-17 | + +Retrieval is 14.9s of that, which is why +[spec 009](../../009-collection-routing/spec.md) is on this feature's critical path +rather than a parallel nicety. + +## What the website side needs to decide + +1. How a verified person is represented (D1) -- the security boundary +2. Whether it calls this service directly from the browser, or proxies server-side. + Proxying makes the token question easier and keeps this service off the public + internet +3. What it renders for each `state`, particularly `nothing_found` -- probably nothing +4. Whether repeat searches are cached at its layer or this one (D3) diff --git a/specs/010-search-page-answers/spec.md b/specs/010-search-page-answers/spec.md new file mode 100644 index 0000000..9cc69eb --- /dev/null +++ b/specs/010-search-page-answers/spec.md @@ -0,0 +1,193 @@ +# Feature Specification: Chatbot Answers in the Website Search Results + +**Feature Branch**: `010-search-page-answers` + +**Created**: 2026-09-17 + +**Status**: Draft. This repo's half only. Three decisions (D1-D3), and one +measured constraint that shapes all of them. + +**Input**: *"a feature we want is for the chatbot responses to be integrated into +the website search page results in a google like style ... this would have to be a +coordinated effort with the website repo ... the functionality should only be +available to proven to be human users."* + +## The constraint that shapes everything + +Measured on 2026-09-17, beta, Release97 bundle, `gpt-4o-mini`, warm process: + +| | | +|---|---| +| Whole answer | **25.4s and 32.2s** on two runs of the same question | +| Retrieval alone | 14.9s of that | +| Answer length | ~4,400 characters | +| Graph construction | 51.5s, once at startup | + +A Google AI overview arrives in roughly one to two seconds. At twenty-five, a panel +on the search results page spins for the entire time a person reads the ordinary +results and leaves. + +This is not a polish problem. It decides whether the feature works, so it is stated +here as a requirement rather than discovered during implementation. + +Two consequences run through the rest of this spec: the answer **must stream**, so +something appears in the first second or two; and the work in +[spec 009](../009-collection-routing/spec.md) is on this feature's critical path, +because retrieval is the largest single component. + +## User Scenarios & Testing + +### User Story 1 - A verified person searches and sees an answer forming (Priority: P1) + +Someone searches reactome.org. Above the ordinary results, an answer begins appearing +within about two seconds and completes in under ten, with citations into Reactome. + +**Why P1**: it is the feature. Everything else is a refinement of it. + +**Independent test**: post a question to the answer endpoint with a valid human +token; assert the first token arrives inside the budget and citations resolve to real +stable IDs. + +**Acceptance** +1. First streamed token within **2s** (FR-005) +2. Complete answer within **10s** (FR-005) +3. Every factual claim carries a citation resolvable at `reactome.org/content/detail/` +4. A question the knowledgebase cannot answer says so, rather than inventing + +### User Story 2 - Someone who has not been verified gets no answer (Priority: P1) + +An unverified visitor, or a script, gets the ordinary search results and no AI panel. +No LLM call is made. + +**Why P1 and not P2**: this is a cost and abuse control, not a feature toggle. Every +search reaching the model is an unbounded bill, and search pages are crawled. + +**Independent test**: post without a token, and with a forged one; assert HTTP 401/403 +and that no LLM call was made. + +**Acceptance** +1. No valid human token → refused before any model call +2. A token is bound to the person, expires, and cannot be replayed from elsewhere +3. Refusal is cheap and does not consume a rate-limit slot for real users + +### User Story 3 - Not every search gets an answer (Priority: P2) + +A search for a single gene name, or a navigational query, returns ordinary results +with no AI panel. Only questions that an answer would actually serve get one. + +**Why P2**: the feature works without it, but the cost does not. + +**Independent test**: a set of queries labelled should-answer / should-not; assert the +classifier's decision matches, and that no LLM answer call happens for the latter. + +### Edge Cases + +- **Retrieval finds nothing**: say so; never fill the gap with general knowledge +- **The model is slow or the upstream is down**: the panel disappears rather than + hanging; the search page must not depend on this service being up +- **The person navigates away mid-stream**: the request is cancelled, not left running +- **The same question twice**: served from cache, not re-answered (see D3) +- **A question that is unsafe or off-topic**: the existing safety checker already + refuses these, and its refusal must not render as an AI panel + +## Requirements + +### Functional Requirements + +- **FR-001**: The service MUST expose an HTTP endpoint accepting a question and + returning an answer with citations. No such endpoint exists today -- `chat-fastapi.py` + serves only the captcha pages and a landing page, and Chainlit owns the conversation + over websockets +- **FR-002**: The response MUST stream, so partial text can render before completion +- **FR-003**: The endpoint MUST refuse any request without a valid proof-of-human + token, before any model call +- **FR-004**: Citations MUST be Reactome stable IDs, so the website can render links + in its own style rather than parsing prose +- **FR-005**: First token within **2s** and complete within **10s**, at p50, for a + question the knowledgebase can answer +- **FR-006**: The service MUST fail invisibly: any error, timeout or refusal returns + a response the website can render as "no panel", never a broken panel +- **FR-007**: Answers MUST be attributable to a Reactome release, so a cached or + stale answer can be identified after a release +- **FR-008**: The endpoint MUST be rate limited per verified person, independently of + the chat UI's existing limits + +### Key Entities + +- **Question**: the search string, plus the release it was asked against +- **Answer**: streamed text, a list of citations (stable ID and display name), and a + completion state -- answered, nothing-found, refused, or failed +- **Human token**: evidence that a person was verified, bound to them, time limited. + Turnstile already does this in `chat-fastapi.py`; what is missing is a form the + website can obtain and present + +## Success Criteria + +### Measurable Outcomes + +- **SC-001**: p50 first token ≤ 2s, p50 complete ≤ 10s, measured on the tracked + question set against a real bundle -- today the whole answer is 25-32s +- **SC-002**: Zero model calls for requests without a valid token, measured by + counting calls under a load of unauthenticated requests +- **SC-003**: The answer sweep stays green: the endpoint and the chat UI give the + same answer to the same question, because they share a graph +- **SC-004**: No search-page request can make the search page itself slower or fail; + verified by taking the service down and confirming the page still renders + +## Decisions + +### D1 -- what proves a person is human? + +Turnstile already exists here, and a bug in it was fixed on 2026-09-17: a deployment +mounting `CLOUDFLARE_SECRET_KEY` as a Docker secret had the captcha silently +disabled, because the middleware read `os.environ` directly while the value came from +`get_secret`. + +What is missing is the handoff. The website needs something to present to this +service. Options: a signed cookie on the shared parent domain (what the chat uses +now); a short-lived token the website mints after its own Turnstile check; or the +website proxying the call and vouching server-side. + +**Open.** It is a security boundary and belongs with whoever owns the website's +session model, not with this repo alone. + +### D2 -- which searches get an answer? + +Not all of them, per User Story 3. The intent classifier already decides what kind +of question it is, and [spec 007](../007-answer-cascade/spec.md) is about that +routing. Reusing it costs nothing extra; inventing a second classifier costs a call. + +**Recommendation**: extend the existing classifier rather than add one. **Open** +pending 007. + +### D3 -- caching + +Search traffic repeats in a way chat traffic does not. A cache keyed by question and +release would cut both latency and cost, and FR-007 exists to make invalidation +possible. + +**Open.** It needs a measurement of repeat rate on real search traffic, which this +repo does not have -- the website does. + +## Scope + +**In (this repo)**: the answer endpoint, streaming, citation shape, token +verification, rate limiting, release attribution, and the latency work needed to +meet FR-005. + +**Out (the website repo)**: the panel, its visual design, where it sits on the page, +how citations render, and the UX of the chat's deeper integration. This spec commits +to a contract, not a look. + +**Out (both, for now)**: the broader "chat integrated into the site's flow and feel" +work. That is a design effort that deserves its own spec once this contract exists, +and it is mostly not this repo's code. + +## Assumptions + +- The website can verify a human before calling this service; this repo verifies the + evidence rather than performing the check +- Beta is the integration target first. The MCP already points at + `beta.reactome.org` as of 2026-09-17 +- The same graph serves the endpoint and the chat UI; two answer paths that could + disagree would be a defect, not a feature From 0474c6bd1e2b508a27bc00ccca84eccbfef45253 Mon Sep 17 00:00:00 2001 From: Adam Wright Date: Thu, 17 Sep 2026 16:50:15 +0000 Subject: [PATCH 2/2] Correct spec 010's latency numbers, which I overstated Three faults, found by attacking yesterday's spec rather than defending it. The headline was wrong. I quoted "25.4s and 32.2s" as the answer latency; that was two runs of one question, and that question sits near the maximum. Measured across all fifteen tracked questions: min 6.7s, p50 15.2s, p90 22.4s, max 31.5s. The figure I built a requirement on was roughly twice the median. The retrieval attribution was wrong in a way that mattered more. I wrote "retrieval is 14.9s of that", combining a sync-path measurement with an end-to-end number from a different run. Retrieval is about 12.5s of a 27s answer on a heavy reactome question and a small fraction of a 12.8s userguide one, so stating it as a single share across all questions was never right. And the conclusion I drew from it does not survive. Retrieval scales with queries x collections: query expansion turns one question into five queries for 2.4s and one LLM call, and each runs against every collection. Five queries and five collections is 12.5s; one query and five collections is about 1.5s. Cutting queries is a lever the same size as cutting collections, and only the second has a spec. Spec 009 is still worth doing; the claim that it alone is on the critical path was mine and it was wrong. The Google "one to two seconds" figure is now marked as the assumption it is rather than presented as something we measured. The conclusion holds anyway: at a p50 of fifteen seconds a search panel is still spinning long after the reader has gone. Also recorded in the contract, because it will mislead whoever optimises this next: the async retrieval path is not faster than the sync one -- 12.5s against 10.9s on the same five queries. The concurrency is not currently buying throughput. Co-Authored-By: Claude Opus 5 --- .../contracts/answer_endpoint.md | 18 +++++--- specs/010-search-page-answers/spec.md | 44 ++++++++++++++----- 2 files changed, 46 insertions(+), 16 deletions(-) diff --git a/specs/010-search-page-answers/contracts/answer_endpoint.md b/specs/010-search-page-answers/contracts/answer_endpoint.md index 94778a7..0e4b231 100644 --- a/specs/010-search-page-answers/contracts/answer_endpoint.md +++ b/specs/010-search-page-answers/contracts/answer_endpoint.md @@ -72,11 +72,19 @@ the same question, that is a defect. | | target | today | |---|---|---| | first token | 2s | n/a -- no streaming endpoint exists | -| complete | 10s | **25.4s and 32.2s** measured 2026-09-17 | - -Retrieval is 14.9s of that, which is why -[spec 009](../../009-collection-routing/spec.md) is on this feature's critical path -rather than a parallel nicety. +| complete | 10s | p50 **15.2s**, p90 **22.4s**, max **31.5s** over the 15 tracked questions | + +Retrieval dominates the heavy questions -- about 12.5s of a 27s answer -- and it +scales with *queries x collections*. Query expansion turns one question into five +queries (2.4s, one LLM call) and each runs against every collection, so cutting +queries is a lever of the same size as cutting collections. Only the latter has a +spec. See [009](../../009-collection-routing/spec.md), and note that it is one lever +rather than the whole of it. + +One measurement worth keeping in view for anyone optimising this: the async +retrieval path is **not faster than the sync one** here -- 12.5s against 10.9s on the +same five queries. The concurrency in `aretrieve_documents` is not currently buying +throughput, so a plan that assumes it will is assuming something unmeasured. ## What the website side needs to decide diff --git a/specs/010-search-page-answers/spec.md b/specs/010-search-page-answers/spec.md index 9cc69eb..a5d91e0 100644 --- a/specs/010-search-page-answers/spec.md +++ b/specs/010-search-page-answers/spec.md @@ -16,24 +16,46 @@ available to proven to be human users."* Measured on 2026-09-17, beta, Release97 bundle, `gpt-4o-mini`, warm process: +Across all fifteen tracked questions, not one question twice: + | | | |---|---| -| Whole answer | **25.4s and 32.2s** on two runs of the same question | -| Retrieval alone | 14.9s of that | -| Answer length | ~4,400 characters | +| Whole answer | min **6.7s**, p50 **15.2s**, p90 **22.4s**, max **31.5s** | +| Retrieval step, heavy reactome question | ~12.5s of a 27s answer | +| Retrieval step, userguide question | a fraction of a 12.8s answer | +| Query expansion | 2.4s, one LLM call, producing 4 variants + the original | | Graph construction | 51.5s, once at startup | -A Google AI overview arrives in roughly one to two seconds. At twenty-five, a panel -on the search results page spins for the entire time a person reads the ordinary -results and leaves. +*An earlier draft of this spec quoted "25.4s and 32.2s" as the headline. That was two +runs of one question, and that question is near the maximum -- roughly twice the p50. +Corrected on 2026-09-17 after measuring the distribution.* + +A Google AI overview is generally reported to arrive in one to two seconds -- an +assumption here, not a measurement of ours. At a p50 of fifteen seconds, a panel on +the search results page is still spinning long after a person has read the ordinary +results, so the conclusion survives the correction even though the number did not. This is not a polish problem. It decides whether the feature works, so it is stated here as a requirement rather than discovered during implementation. -Two consequences run through the rest of this spec: the answer **must stream**, so -something appears in the first second or two; and the work in -[spec 009](../009-collection-routing/spec.md) is on this feature's critical path, -because retrieval is the largest single component. +Two consequences run through the rest of this spec. + +The answer **must stream**, so something appears in the first second or two. + +And retrieval is the largest single component for the questions that matter here -- +but **collection routing is not the only lever on it, and may not be the biggest**. +Retrieval cost scales with *queries x collections*. Query expansion turns one +question into **five** queries at a cost of 2.4s, and every one of them is run +against every collection. Cutting five collections to two saves about as much as +cutting five queries to two, and only the first has a spec. Measured 2026-09-17: + +| | | +|---|---| +| 5 queries x 5 collections, served async path | 12.5s | +| 1 query x 5 collections, served async path | ~1.5s | + +[Spec 009](../009-collection-routing/spec.md) remains worth doing. The claim that it +alone is on the critical path does not survive the measurement. ## User Scenarios & Testing @@ -126,7 +148,7 @@ classifier's decision matches, and that no LLM answer call happens for the latte ### Measurable Outcomes - **SC-001**: p50 first token ≤ 2s, p50 complete ≤ 10s, measured on the tracked - question set against a real bundle -- today the whole answer is 25-32s + question set against a real bundle -- today p50 is 15.2s and p90 22.4s - **SC-002**: Zero model calls for requests without a valid token, measured by counting calls under a load of unauthenticated requests - **SC-003**: The answer sweep stays green: the endpoint and the chat UI give the