Skip to content

Spec 010: plan, research and 24 tasks — build the endpoint slow, unblock the website repo - #232

Merged
adamjohnwright merged 2 commits into
mainfrom
spec/010-plan-and-tasks
Sep 17, 2026
Merged

adamjohnwright merged 2 commits into
mainfrom
spec/010-plan-and-tasks

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Spec 010 had a spec and a contract and no plan. This adds plan.md, research.md, quickstart.md and 24 tasks — and takes a position on sequencing.

The sequencing call

Build the endpoint first, and slow. The website session is blocked on nothing but this service's absence. An endpoint that answers correctly in fifteen seconds unblocks their panel and their token handshake; an endpoint that doesn't exist blocks them completely.

So FR-005's 2s/10s budget is explicitly not a gate on the first increment. MVP is Phases 1–4: correct, verified, streaming, at roughly today's p50. Latency is Phase 5 and runs alongside the handover.

Three things from reading the code rather than assuming

AgentGraph exposes ainvoke and nothing else. Chainlit's streaming comes from a Chainlit callback handler, not any graph-level API — so there is no surface an HTTP handler can use. The plan takes LangGraph's astream_events, names a queue-fed callback as the fallback, and requires a test that tokens actually arrive. A surface that completes without streaming would pass a 200 check, which is exactly how the userguide shell fooled two sessions today.

The FastAPI app doesn't hold a graph at all — chat-chainlit.py constructs it. Building a second would pay 51.5s of startup twice and let the panel and the chat answer differently, which SC-003 forbids. One instance, shared.

Citations come from retrieval, not prose. The answer prompt emits <a href> anchors inline, and parsing those out of a token stream would couple the wire contract to prompt wording — which changed twice this week. Retrieved documents already carry st_id. The cost is stated rather than hidden: citations become what retrieval found rather than what the answer used.

Where two principles meet

FR-006 says fail invisibly; Principle IV says fail loudly. Not in conflict, but the line is written down:

  • Misconfiguration stops the process — a missing verifying key means the endpoint refuses to start, because one that accepts everything would test clean
  • A runtime failure answering one question returns state: failed and logs, because a search page must not break

Also noted

verify_captcha_middleware intercepts every path under CHAINLIT_URI, so a new route needs it to pass /chat/api/ through — with a characterization test pinning the middleware's current behaviour first. That middleware is where today's silent-captcha-bypass bug lived, so adding a route beside it warrants the care.

Docs only, no code. 333 tests still green.

🤖 Generated with Claude Code

adamjohnwright and others added 2 commits September 17, 2026 17:48
Spec 010 had a spec and a contract and no plan. This adds plan, research,
quickstart and 24 tasks, and takes a position on sequencing: build the endpoint
first and slow, rather than making the website repo wait for latency work.
Their side is blocked on nothing but this service's absence. An endpoint that
answers correctly in fifteen seconds unblocks them; one that does not exist
does not.

Three things came out of reading the code rather than assuming.

AgentGraph exposes ainvoke and nothing else. Chainlit's streaming comes from a
Chainlit callback handler, not from any graph-level API, so there is no
surface an HTTP handler can use. The plan takes LangGraph's astream_events with
a queue-fed callback named as the fallback, and requires a test that tokens
actually arrive -- a surface that completes without streaming would pass a 200
check, which is precisely how the userguide shell fooled two sessions today.

The FastAPI app does not hold a graph at all; chat-chainlit.py constructs it.
Building a second would pay 51.5s of startup twice and let the panel and the
chat answer differently, which SC-003 forbids.

Citations come from retrieval, not from prose. The answer prompt emits <a href>
anchors inline, and parsing those out of a token stream would couple the wire
contract to prompt wording -- which changed twice this week. Retrieved
documents already carry st_id, so citations are emitted from them. The cost is
stated rather than hidden: citations become what retrieval found rather than
what the answer used.

The constitution check records where FR-006 and Principle IV meet. Fail
invisibly and fail loudly are not in conflict but the line matters:
misconfiguration stops the process, a runtime failure answering one question
returns state failed. A missing verifying key is the first kind, because an
endpoint that accepts everything would test clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 4a77776 into main Sep 17, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the spec/010-plan-and-tasks branch September 17, 2026 17:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant