Skip to content

Rate limit the answer endpoint, and stop paying for a search nobody reads - #237

Merged
adamjohnwright merged 1 commit into
mainfrom
feat/answer-rate-limit
Sep 18, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
feat/answer-rate-limit

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

The two items left open after #236 — FR-008 rate limiting (T017) and SC-003 (T014) — plus a third defect found while testing them.

FR-008: a backstop limit (T017)

30 requests per 10 minutes per caller, refused with the same done shape as any other refusal, so the page renders no panel rather than a broken one. The website enforces the real budget upstream; this exists so a leaked or shared token cannot run up an unbounded bill against a service whose every answer costs six model calls.

Keyed on sub, then jti, then a hash of the token. human_token's claims are D1 and still unsettled with the website, so sub may never arrive — this starts counting real people the moment it does, and uses the token as a proxy until then. Hashed because the key outlives the request and a bearer token is a credential.

Stale keys are swept once per window, or the dict grows by one entry per visitor forever. Both the sweep and its absence are mutation-tested: breaking the sweep in either direction fails a test.

The endpoint was buying a web search and throwing it away

postprocess runs a Tavily web search after the answer and writes it to additional_content — and astream_answer has no event that could carry that. So every search-page answer paid for a web search, discarded the result, and delayed done by its duration. The chat UI renders those results; this surface has nowhere to put them. Now enable_postprocess=False.

T014 could not be done as written, and the spec was wrong

SC-003 asked to assert that the endpoint and the chat UI "give the same answer". They do not — and neither does either surface give the same answer as itself.

Measured 2026-09-18, one graph, temperature 0:

comparison similarity
endpoint run A vs endpoint run B (same surface) 0.331
endpoint vs chat 0.356

Each surface differs from itself as much as it differs from the other. Asserting equality would be asserting that a model is deterministic.

So I corrected the criterion rather than weakening the test. SC-003 is now configuration equivalence — same profile, same shared graph, no difference that can reach the answer — pinned by tests that need no model:

  • the endpoint's PROFILE equals ProfileName.React_to_Me.lower(), so renaming the profile can no longer leave the endpoint pointing at a key that does not exist (a missing key makes astream_answer yield failed with no other signal)
  • both surfaces take their graph from the shared registry
  • the one deliberate difference, enable_postprocess, is pinned as safe because postprocess runs after the answer — with a test that fails if it ever starts editing the answer instead

The live comparison is kept, opt-in behind a marker, asserting both paths produce a substantive answer rather than identical prose. It was run for real, not just written.

Also found: retrieval is not reproducible

Three runs of one question returned 12 citations each but shared only 4, union 19, Jaccard 0.26 — query expansion is itself a model call, so each run expands differently and retrieves different documents.

A reader who reloads the panel sees different sources, and FR-007's cache invalidation assumes an answer is a stable artifact of a release. Recorded in the spec and left open as T020a; not a handover blocker and deliberately not decided here.

CI-equivalent locally: ruff, format, mypy (128 files), full suite.

🤖 Generated with Claude Code

Two open items from spec 010, plus a third found while testing them.

**FR-008 (T017).** A backstop behind the website's own budget: 30 requests per
10 minutes, refused with the same `done` shape as any other refusal so the page
renders no panel. Keyed on `sub` then `jti`, so it starts counting real people
the moment D1 settles, and on a hash of the token until then -- hashed because
the key outlives the request and a bearer token is a credential. Stale keys are
swept, or the dict grows by one entry per visitor forever.

**The endpoint was buying a web search and throwing it away.** `postprocess`
runs a Tavily search after the answer, and `astream_answer` has no event that
could carry the result -- so every search-page answer paid for one, discarded
it, and delayed `done` by its duration. The chat UI renders those results; this
surface has nowhere to put them. Now `enable_postprocess=False`.

**T014 could not be done as written.** SC-003 asked to assert the endpoint and
the chat UI give the same answer. Measured at temperature 0 on one graph: the
same question through the *same* surface twice scored 0.331 similarity, and
endpoint-versus-chat scored 0.356 -- each surface differs from itself as much as
from the other. Asserting equality would assert that a model is deterministic.

So the spec's criterion was corrected, not the test weakened: SC-003 is now
configuration equivalence, pinned by tests that need no model, plus the sweep,
which already matches patterns rather than literals.

Also recorded: retrieval is not reproducible either. Three runs of one question
returned 12 citations each but shared only 4, union 19, Jaccard 0.26, because
query expansion is a model call too. That bears on what FR-007 can cache, and is
left open as T020a rather than decided here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit 32ca9b1 into main Sep 18, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the feat/answer-rate-limit branch September 18, 2026 02:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant