Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
84 changes: 84 additions & 0 deletions specs/010-search-page-answers/research.md
Original file line number Diff line number Diff line change
Expand Up @@ -245,3 +245,87 @@ preprocessing, so a sequencing confound would have shown up there and did not.
calls.
3. **The answer model, 0.84s** -- not worth attention on these numbers, which
is exactly why the discrepancy above matters.

## T020c -- does a search-page question need all four preprocessing calls?

Measured 2026-09-20. The answer is yes for three of them, and "not for the
reason you would guess" for the fourth.

### The rephraser is not a no-op without history

It exists to resolve a follow-up against chat history, and the search page
sends one question on a fresh thread. The obvious conclusion is that it does
nothing there. Measured over the fifteen tracked questions with empty
history, it **changed 14 of 15**:

- "Which release of Reactome is this?" -> "Which version of Reactome is
currently available?"
- "What does CDK5 phosphorylate in Alzheimer disease?" -> "What are the
substrates that CDK5 phosphorylates in the context of..."

That is query normalisation rather than follow-up resolution: it feeds
retrieval and intent classification with different text than it was given.

**"Changed" is not "improved", and the two facts here point opposite ways.**
It rewrites 14 of 15 questions, *and* the sweep passes without it -- so for
the tracked set those rewrites are not load-bearing. What the measurement
supports is that removing it is a retrieval-*input* change rather than a
plumbing one, and therefore needs measuring on something other than the
thirteen questions that already pass either way. It does not support the
claim that the rewrites are valuable.

### With it bypassed, the sweep still passes -- and latency barely moves

`answer-sweep` was 13/13 with the rephraser replaced by a pass-through
(precondition asserted: bypassed 13 times).

Latency, both arms **interleaved in one window**, n=25 each:

| | first token p50 | min | max |
|---|---|---|---|
| expansion off | 2.94s | 2.35s | 5.94s |
| expansion off + rephrase off | **2.83s** | 2.04s | 8.12s |

About **0.11s**, which is the round arithmetic below rather than the ~0.8s the
call's own duration suggests.

A first version of this compared 2.97s against a 2.73s taken in an earlier
run and reported "no better, slightly worse" -- the cross-window comparison
that the T022 review had just corrected, repeated here a few hours later. The
direction was wrong; the conclusion that the saving is ~0.1s rather than ~0.8s
was not, and is now better supported.

### Why removing a 0.88s call saved 0.07s

The arithmetic that predicted ~0.8s was wrong, and the error is worth
recording because it is easy to repeat.

Preprocessing runs two rounds. Round one is `max(rephrase, language)` and
round two is `max(safety, intent)`. Median call times are rephrase 0.88s,
language 0.81s, safety 0.70s, intent 0.93s. So round one costs 0.88s with the
rephraser and **0.81s without it** -- language detection is nearly as slow, and
it was hiding behind the rephraser the whole time.

Bypassing the call does not remove the round. The saving needs the rounds
**restructured**: with no rephrasing, `rephrased_input` is just `user_input`,
so safety and intent no longer depend on round one and all three calls can run
together -- 0.93s instead of 1.74s. That is the ~0.8s, and it requires a code
change this measurement did not make.

### So, for FR-005a

**2s to first token is not reachable by removing these two calls.** Expansion
off gets p50 to 2.94s and adding rephrase-off reaches 2.83s -- the best single
observation dipped to 2.04s, and the median did not. Restructuring preprocessing into one round is worth
about 0.8s more on paper, which would put p50 near 2s -- on paper, and against
a tail that neither change addresses.

### What this does not establish

Both levers pass the same thirteen questions, and that is now being asked to
carry a lot. Each alone keeps the sweep green; together they keep it green;
none of that shows recall is unaffected, and removing *two* recall mechanisms
on one thin evidence base compounds the risk rather than adding to the
confidence. The rephraser rewriting 14 of 15 questions is the concrete reason
to be careful: whatever those rewrites are worth, the sweep is not what
measures it.
2 changes: 1 addition & 1 deletion specs/010-search-page-answers/tasks.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ state.
- [x] T019 Measure first-token and completion separately across the tracked questions; publish the distribution, not one question. Measured 2026-09-18, two runs each: first token p50 9.6s / p90 12.2s (n=26), completion p50 10.4s / p90 18.1s (n=30). The earlier "36.1s to first token" came from one question and does not reproduce (PR #238)
- [x] T019a Run preprocessing in two rounds instead of four sequential calls in src/agent/profiles/react_to_me.py; the base class already overlapped, and this override discarded it (PR #238)
- [x] T020 Measure query expansion, and make the count configurable. **The task's premise was wrong**: the cost is the expansion *call*, not the number of variants it returns. Measured over six tracked questions -- 4 alternates 1.27s expand + 1.22s retrieve; 2 alternates 1.26s + 0.65s; 0 alternates 0.00s + 0.31s. Trimming the count saves only fan-out; the cost disappears only at zero, where `answer-sweep` was 13/13 in 79s against ~150s. `QUERY_EXPANSION_ALTERNATES` sets it and **the default is unchanged**: thirteen questions show those answers do not need expansion, not that recall is unaffected in general
- [ ] T020c Establish whether a search-page question needs all four preprocessing calls. The sequential half of this is answered and done (T019a): they run in two rounds and cost 2.6s at the median, not the ~16s recorded here, which never reproduced
- [x] T020c Established 2026-09-20. **The rephraser is not a no-op without history** -- it changed 14 of 15 tracked questions, normalising queries rather than resolving follow-ups, so removing it is a retrieval-quality change. Bypassed, the sweep is 13/13 and first token p50 goes 2.94s to 2.83s (interleaved, n=25): about 0.11s. Removing a 0.88s call saved 0.07s because round one is `max(rephrase, language)` and language detection is 0.81s -- it was hiding behind the rephraser. The ~0.8s needs the **rounds restructured** so safety and intent stop waiting, which is a code change this did not make. 2s is not reachable by removing these two calls
- [ ] T021 Land spec 009 collection routing and re-measure. **Still open — I marked this done on 2026-09-18 and was wrong.** What landed is *source* routing (`resolve_active_sources` picks reactome / userguide / live). Collection routing is selecting among the five collections *within* the reactome bundle, and it is not implemented: `QueryIntent` has no `collections` field, `resolve_collections` does not exist, and `retrieve_documents` still loops over every collection
- [x] T022 Re-assessed 2026-09-20 on the served path, arms interleaved, n=25 each, with the query count asserted rather than assumed. **10s complete is met** (8.90s p50, 7.16s with expansion off); **2s to first token is not** (5.57s p50, 2.73s with expansion off), so FR-005a stays a target with a named blocker. The blocker changed: retrieval is the largest block at 3.15s and the answer model is 0.84s against a previously recorded 6.1s, recorded as unreconciled. The tail is the finding -- default first token ranges 2.26s to 11.44s, and expansion barely moves the best case while halving the worst, so a future target belongs at a high percentile
- [x] T023 Propagate a real failure signal out of the live path. `answer_from_live_services` takes an optional `LiveReport` and records a tool exception **before** stringifying it into the model's context -- the last point at which it is still a fact rather than a paraphrase. The graph carries it out as `live_tool_failed`, and `answer_sweep` retries on that instead of matching prose. `TRANSIENT` and `_looks_transient` are gone, with a test asserting they do not come back: if prose matching returns it will return as a list of markers. The report is optional, so none of the twelve existing call sites changed. **The no-tools fallback deliberately does not set it**: `get_mcp_tools` remembers a failed start, so a retry could never succeed and marking it transient would buy an attempt guaranteed to fail. Pinned by a test, because the absence of the flag there reads like an oversight
Expand Down
Loading