Repository navigation
Spec 010: chatbot answers in the website search results — this repo's half - #230
Merged
Merged
Conversation
Adam wants chatbot answers in the website search results, Google-overview style, for verified humans only, coordinated with the website repo. The thing worth knowing before anyone designs anything: a whole answer takes 25.4s and 32.2s on two runs of the same question, measured today on beta against Release97. Retrieval is 14.9s of it. A Google AI overview arrives in one to two seconds. At twenty-five a panel spins for the entire time someone reads the ordinary results and leaves. So the spec states the budget as a requirement -- first token in 2s, complete in 10s -- rather than leaving it to be discovered in implementation. Two things follow: the answer has to stream, and spec 009 is on this feature's critical path rather than a parallel nicety, because retrieval is the largest single component. Three things this repo does not have: an answer endpoint at all (chat-fastapi.py serves the captcha pages and a landing page; Chainlit owns the conversation over websockets), a way for the website to present proof that it verified a person, and any rule about which searches deserve an answer. The last matters because search pages are crawled and every crawled search reaching the model is an unbounded bill. Turnstile already exists here, and the bug fixed this morning is directly relevant: a deployment mounting CLOUDFLARE_SECRET_KEY as a Docker secret had the captcha silently disabled. That is the control this feature depends on. contracts/answer_endpoint.md is the part the website session can start from. It commits to a wire shape and, more importantly, to properties: refuse before any model call, fail invisibly so the search page never breaks, citations as stable IDs so the site styles its own links, and one graph behind both the endpoint and the chat so they cannot disagree. Scope is this repo only. The panel, its design, and the wider "chat integrated into the site's feel" work are the website's, and deserve their own spec once this contract exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three faults, found by attacking yesterday's spec rather than defending it. The headline was wrong. I quoted "25.4s and 32.2s" as the answer latency; that was two runs of one question, and that question sits near the maximum. Measured across all fifteen tracked questions: min 6.7s, p50 15.2s, p90 22.4s, max 31.5s. The figure I built a requirement on was roughly twice the median. The retrieval attribution was wrong in a way that mattered more. I wrote "retrieval is 14.9s of that", combining a sync-path measurement with an end-to-end number from a different run. Retrieval is about 12.5s of a 27s answer on a heavy reactome question and a small fraction of a 12.8s userguide one, so stating it as a single share across all questions was never right. And the conclusion I drew from it does not survive. Retrieval scales with queries x collections: query expansion turns one question into five queries for 2.4s and one LLM call, and each runs against every collection. Five queries and five collections is 12.5s; one query and five collections is about 1.5s. Cutting queries is a lever the same size as cutting collections, and only the second has a spec. Spec 009 is still worth doing; the claim that it alone is on the critical path was mine and it was wrong. The Google "one to two seconds" figure is now marked as the assumption it is rather than presented as something we measured. The conclusion holds anyway: at a p50 of fifteen seconds a search panel is still spinning long after the reader has gone. Also recorded in the contract, because it will mislead whoever optimises this next: the async retrieval path is not faster than the sync one -- 12.5s against 10.9s on the same five queries. The concurrency is not currently buying throughput. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The chatbot side of putting answers into the website search page, Google-overview style, for verified humans only. Spec and contract; no code.
The number that decides the design
Measured today on beta against Release97, warm process:
A Google AI overview arrives in one to two seconds. At twenty-five, a panel on the search results page spins for the entire time someone reads the ordinary results and leaves.
That isn't a polish problem, so the spec states it as a requirement — first token in 2s, complete in 10s — rather than leaving it to be discovered mid-implementation. Two things follow:
Three things this repo doesn't have yet
No answer endpoint at all.
chat-fastapi.pyserves the captcha pages and a landing page; Chainlit owns the conversation over websockets. A search panel needs a plain streaming endpoint returning text plus citations.No way for the website to prove it verified a person. Turnstile exists here — and this morning's fix is directly relevant, since a deployment mounting
CLOUDFLARE_SECRET_KEYas a Docker secret had the captcha silently disabled. That's the control this feature depends on. What's missing is the handoff, which is a security boundary and belongs with whoever owns the website's session model (D1, open).No rule about which searches deserve an answer. This one is a cost control, not a toggle: search pages get crawled, and every crawled search reaching the model is an unbounded bill. The intent classifier already decides what kind of question it is, so extending it costs nothing where a second classifier costs a call (D2, pending spec 007).
What the website session can start from
contracts/answer_endpoint.mdcommits to a wire shape, and more importantly to properties:Scope
This repo only: the endpoint, streaming, citations, token verification, rate limiting, release attribution, and the latency work to meet the budget.
The panel, its design, and the wider "chat integrated into the site's flow and feel" work are the website's, and deserve their own spec once this contract exists. This one commits to a contract, not a look.
Created with
create-new-feature.shrather than by hand —.specify/feature.jsonpoints at it, so the Spec Kit stages work.🤖 Generated with Claude Code