Skip to content

Spec 011: summarise analysis results - #252

Merged
adamjohnwright merged 4 commits into
mainfrom
011-summarise-analysis-results
Sep 18, 2026
Merged

adamjohnwright merged 4 commits into
mainfrom
011-summarise-analysis-results

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Spec, plan and tasks only — no implementation. Opening this because the work has been sitting on an unmerged branch, invisible to anyone reading the repo.

Adam asked for endpoints that summarise Reactome analysis results. I researched the analysis surface first, and two findings shaped the design.

Results are token-addressed

GET /token/{token} returns a completed AnalysisResult. So a summary takes a token, not a gene list — we summarise an analysis the user already ran and never run one, which is what Principle V requires anyway.

Three types come from the type enum — OVERREPRESENTATION, EXPRESSION, SPECIES_COMPARISON — with ReactomeGSA a fourth family, flagged by gsaMethod/gsaToken and served from a different host. GSA is recognised and declined in this increment.

Results are deleted on a new release

The API defines exactly two documented errors:

code meaning
404 no result corresponds to the token
410 result deleted due to a new data release

That makes the release part of the storage key rather than decoration — a stored summary can otherwise outlive the result it describes. And gone must be a distinct outcome from not_found: one has an action attached ("run it again"), the other is a dead end.

A third code the API does not document: a malformed token returns 500, measured against beta. Without that, the implementation would have handled the two documented codes and let a malformed token become a failed state or a retry loop.

The privacy trap is not the gene list

Adam's decisions — opt-in, a real choice of what is shared, evidence a person is present, stability, transparency — are requirements FR-011 to FR-015.

Three fields in the aggregate result are user-supplied free text: summary.fileName, summary.sampleName, expression.columnNames. A tier defined as "don't send the identifiers" would pass smith_lab_unpublished_2026.txt or Patient_001_tumour straight through. So the no-disclosure tier is an allow-list of named fields, which is wrong only by omission — a missing sentence rather than a disclosure.

Usefully, that tier answers "what does my result say" completely. Only "which of my identifiers failed" needs the disclosing option, so the choice offered to users is real and explainable.

Stability by storage, not by pretending

FR-014 wants the same token to yield the same summary. Generation cannot provide that — seeded runs measured 0.47 and 0.14 similarity — so summaries are stored against (token, release, tier) and reused, and cached on the start event says which. The interface must not imply the generator is deterministic.

One blocker, and it is not mine

FR-013 requires evidence a person is present. That is stricter than the caller token, which by spec 010's D1 asserts service identity and deliberately says nothing about humanity. The website has a Turnstile gate now and we are working out how presence is asserted. Phase 8 is marked blocked and blocks nothing else — every story can be built and tested behind a refusing gate.

37 tasks. MVP is Phases 1–3.

🤖 Generated with Claude Code

adamjohnwright and others added 4 commits September 18, 2026 18:39
Researched first, specified second. The analysis surface, from beta's Analysis
Service v3 API and the user guide bundle already indexed here:

- results are addressed by a token (GET /token/{token}), so a summary summarises
  an analysis the user already ran. It never runs one, which is what Principle V
  requires anyway
- the type enum is OVERREPRESENTATION | EXPRESSION | SPECIES_COMPARISON, with
  ReactomeGSA a fourth family flagged by gsaMethod/gsaToken and served elsewhere
- a result carries per-pathway entity statistics -- found, total, ratio, pValue,
  fdr, exp[] -- plus resourceSummary, unmatched identifier counts and warnings

Four user stories, in the order a reader actually needs them: what does my result
say; why were my identifiers not found; what do these numbers mean; and the
readings specific to expression and species comparison.

The requirements that matter are the ones about honesty: derive every number from
the result, distinguish p-value from FDR whenever calling something significant,
and say plainly when a result is weak or empty rather than presenting the lowest
p-value as a finding. A summary that narrates significance that is not there is
worse than no summary on a scientific resource.

Two questions are left for Adam rather than defaulted: whether result contents,
which include a user's own submitted identifiers, may go to a third-party model
provider; and whether a summary of a fixed result must be stable, given answers
are measured non-reproducible. Both change scope, and neither has a safe default.

ReactomeGSA is deferred in Assumptions -- separate service, separate result
shape, and including it would double the first increment.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Phase 0 research measured against beta's Analysis Service rather than recalled,
and two findings changed the design.

**Results are deleted on a new release.** The API defines exactly two errors: 404
"no result corresponds to the token" and 410 "result deleted due to a new data
release". So a stored summary can outlive the result it describes, which makes the
release part of the storage key rather than decoration -- the same invalidation
signal spec 010 already publishes. It also means `gone` must be a distinct outcome
from `not_found`: one is a dead end, the other has an action attached.

**The disclosure tier is an allow-list, and the trap is not the gene list.**
Three fields in the aggregate result are user-supplied free text -- `fileName`,
`sampleName` and `expression.columnNames` -- so "don't send the identifiers" would
pass a lab's unpublished filename or a patient sample label straight through. The
aggregate tier is defined by naming what may be sent, which is wrong only by
omission.

Adam's two decisions are settled requirements now: opt-in per request, a real
choice of what is shared with a useful option that discloses nothing, evidence a
person is present, stability by storing the summary, and transparency that it was
generated. Stability comes from storage because generation cannot provide it --
seeded runs still scored 0.47 and 0.14 similarity when measured here.

Human presence is the open blocker: the caller token asserts service identity by
D1 of spec 010 and cannot carry it. The likely shape is a claim minted after the
Turnstile check the website already performs, and that is theirs to agree.

Also recorded: beta has no durable store, so first-increment summaries live in
process memory and are lost on deploy -- honest only because the transparency
requirement already tells a user a summary can be regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
37 tasks. Three properties get tests before their implementation, because each is
a claim about something not happening and none fails visibly: what is never sent,
that a dead token is distinguished from one deleted by a release, and that the
same token returns the same text.

The disclosure test asserts on the recorded outbound request rather than on the
summary, because reading the output and seeing nothing alarming is not evidence
that nothing was sent.

Phase 8 -- human presence -- is marked blocked on the website repo rather than
assumed. It blocks nothing else: every story can be built and tested behind a
refusing gate, which is the point of keeping it explicit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Measured against beta: a well-formed but unknown token returns 404 as the
OpenAPI says, but a malformed one returns 500, which the spec does not mention.

FR-009 already calls a malformed token a normal negative outcome. Without this
the implementation would have handled the two documented codes and let a
malformed token become a failed state, or a retry loop against a service that
will answer the same way every time.

Found while adversarially reviewing PR #250, by probing the endpoint the contract
depends on rather than trusting its documentation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@adamjohnwright
adamjohnwright merged commit d45a51c into main Sep 18, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the 011-summarise-analysis-results branch September 18, 2026 18:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant