Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
70 changes: 63 additions & 7 deletions specs/011-summarise-analysis-results/contracts/summary_endpoint.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,8 @@
# Contract: the analysis summary endpoint

Draft. Not implemented, and not yet agreed with the website repo. The parts
marked **open** need their agreement before anything is built against them.
Not implemented. **Agreed with the website repo on 2026-09-19** -- the request
shape, the refusal shape and how human presence is asserted are settled, and
they have the presence claims built on a branch. Nothing here is open.

## Request

Expand All @@ -14,15 +15,64 @@ Content-Type: application/json
"disclosure": "aggregate" }
```

**The analysis token is passed straight through, never validated by the
caller.** This service already distinguishes 404, 410 and the undocumented 500
a malformed token returns (research D2). A second validator on the website
would be a second thing that believes it knows what a good token looks like,
and two validators disagree eventually.

`disclosure` is `aggregate` or `identifiers`, and is **required** — there is no
default, because a default is not a choice. `aggregate` never transmits the user's
identifiers, filenames, sample names or expression column labels.

**Open**: how the caller demonstrates a person is present. Spec 010's D1 settled
that `caller_token` asserts service identity and says nothing about humanity, so
it cannot carry this on its own. The likely shape is an additional claim minted
after the Turnstile check the website already performs on the chat, but that is
theirs to agree.
### Human presence

`caller_token` carries three additional claims, minted only when the website's
Turnstile-backed identity cookie validated on that request:

| claim | meaning |
|---|---|
| `human` | `true`. **Absent otherwise, never `false`** -- a missing claim and a failed check are indistinguishable here |
| `human_iat` | epoch seconds, when the challenge was solved |
| `human_sub` | the cookie's random 16-byte subject, for per-person rate limiting |

**Freshness is 30 minutes and the bound is inclusive**: `now - human_iat <= 1800`,
so a challenge solved exactly 1800s ago is accepted and 1801s is not.

Stated in whole seconds on purpose. The bound was first agreed as "1800.000
accepted, 1800.001 refused", which is not a distinction this claim can carry:
`human_iat` is epoch **seconds**, and it is derived as the cookie's expiry
minus a constant TTL, so it arrives already rounded to a second. A sub-second
edge would be a boundary neither side can actually be on, tested against a
clock finer than the value. Effective precision is one second, and a challenge
solved 1800.4s ago presents as 1800 and is accepted.

Both sides enforce it --
the website refuses to mint past it, this service refuses to accept past it --
so neither is a single point of failure, and if it is ever changed it is
changed in both places in one change.

**`human_sub`, not `sub`.** `sub` already means something here: an opaque
*per-visit* id, which `identity_of` uses to key the answer endpoint's backstop
limiter. The cookie subject is also 16 bytes but is per-*browser* and lives as
long as the cookie. Putting it in `sub` would silently change that limiter from
"this visit" to "this browser", on an endpoint neither repo is otherwise
touching. A separate claim also makes it explicit that a persistent
pseudonymous identifier is being held -- hashed, in memory only.

**And the model of `sub` above was itself wrong.** The website corrected it on
2026-09-19: their `callerSubject()` prefers the identity cookie's subject
whenever the reader has passed a challenge, falling back to the per-visit value
only when they have not. So `sub` has *already* been browser-scoped for every
gated reader, and the answer endpoint's backstop limiter has been counting
across visits since the gate shipped. That is the stronger throttle and it is
kept; the source comments describing it as per-visit are fixed.

A consequence to write down: while `callerSubject()` prefers the verified
identity, `sub` and `human_sub` carry the **same value** whenever both are
present. They are still separate claims, because they mean different things and
would diverge the moment that preference changed -- and because agreement
between them is not a signal anything should test.

## Response: Server-Sent Events

Expand All @@ -47,6 +97,12 @@ data: {"state": "summarised", "seconds": 6.2}
`failed`. Anything but `summarised` means render no summary. Always HTTP 200 —
never an error code, so the analysis page cannot be broken by this service.

**No prose source list.** Citations arrive as `citation` events, as on
`/api/answer`, and the summary text must not end with a list of them --
`SourcesSectionStripper` exists because that duplicate cost the website a
pattern it could not write correctly. This endpoint's prompt is its own, so
the right fix here is not to ask for one in the first place.

`cached` on `start` says whether this text was generated now or reused. It exists
because the interface must not imply determinism it does not have: a reader who
regenerates may get different wording, and `cached: false` is when that happens.
Expand Down
13 changes: 10 additions & 3 deletions specs/011-summarise-analysis-results/research.md
Original file line number Diff line number Diff line change
Expand Up @@ -138,10 +138,17 @@ on the request:
| claim | meaning |
|---|---|
| `human` | `true`, set only when a valid, unexpired identity cookie was presented. **Absent otherwise, never `false`**, so a missing claim and a failed check are indistinguishable to us |
| `human_iat` | when the challenge was solved, epoch seconds. Derived from the cookie's expiry minus their identity TTL; no cookie format change |
| `subject` | the cookie's random 16-byte identifier, for per-identity rate limiting. Carries nothing about the person |
| `human_iat` | when the challenge was solved, epoch **seconds**. Derived from the cookie's expiry minus their identity TTL, so it arrives rounded to a second -- the bound is whole-second, and a sub-second edge is not a state this claim can represent |
| `human_sub` | the cookie's random 16-byte identifier, for per-identity rate limiting. Carries nothing about the person |

**Freshness is 30 minutes, enforced at both ends.** We refuse a `human_iat`
**Named `human_sub`, not `sub`.** `sub` is already an opaque *per-visit* id
that `identity_of` uses for the answer endpoint's backstop limiter. The cookie
subject is per-*browser* and lives as long as the cookie, so reusing the claim
would change that limiter's meaning on an endpoint neither repo is touching --
a behaviour change arriving through a rename. Caught before either side built
to it.

**Freshness is 30 minutes, inclusive in whole seconds (`now - human_iat <= 1800`), enforced at both ends.** We refuse a `human_iat`
older than that, and they refuse to mint the claim past it, so neither side is
a single point of failure. Long enough that reading a result, choosing a
disclosure tier and requesting a summary is never re-challenged; short enough
Expand Down
13 changes: 11 additions & 2 deletions src/util/caller_token.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,17 @@

What it does assert is **caller identity**: this request came from the Reactome
website's server, for one visit. They mint server-side per request, EdDSA, with
`iss`, `aud`, `iat`, `exp` at +120s, and `sub` -- an opaque per-visit id that is
128 random bits, not derived from anything about the reader. Abuse control is
`iss`, `aud`, `iat`, `exp` at +120s, and `sub` -- 128 random bits, not derived
from anything about the reader.

**`sub` is browser-scoped for a gated reader, not per-visit**, and this said
otherwise until 2026-09-19. Their `callerSubject()` prefers the Turnstile
identity cookie's subject whenever the reader has passed a challenge, and falls
back to a per-visit cookie only when they have not. So the backstop limit below
has been keyed on an identifier lasting as long as that cookie -- the stronger
throttle, and the behaviour in production since the gate shipped. It is kept
deliberately; what was wrong was this description of it, which named the
fallback as though it were the only case. Abuse control is
theirs: the panel is opt-in behind a click, so a crawled search never reaches a
model, and their proxy rate limits by address.

Expand Down
12 changes: 8 additions & 4 deletions src/util/rate_limit.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,10 +30,14 @@ def _positive_int(name: str, default: int) -> int:
def identity_of(claims: dict[str, object], token: str) -> str:
"""Who to count against.

The token's claims are D1 and not yet settled with the website, so `sub` may
never arrive. `sub` then `jti` are used when present, so this starts keying on
a real person the moment D1 lands; until then a hash of the token itself is
the best available proxy -- one issuance, short lived, one person.
`sub` then `jti` when present, else a hash of the token itself.

D1 is settled now: `sub` does arrive, and for a reader who has passed the
Turnstile challenge it is the identity cookie's subject -- browser-scoped
for the life of that cookie, not per-visit. So this counts a person across
visits rather than within one, which is the stronger backstop and is what
has been running since the gate shipped. Worth knowing before reasoning
about what a burst here means.

Hashed, never raw: this lands in a dict that lives as long as the process, and
a bearer token is a credential.
Expand Down
Loading