Skip to content

Spec 002: choosing the default answering model; close out spec 001 - #188

Merged
adamjohnwright merged 1 commit into
mainfrom
spec/default-llm-choice
Sep 9, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
spec/default-llm-choice

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Two pieces of design tracking that were owed.

Spec 002 — the gpt-5.6-luna decision

#186 made the model usable. This is about whether it becomes the default, which is not a dependency bump — it changes what every user reads, trades determinism for capability, doubles latency, moves cost in an unmeasured direction, and the same question returns for chat-alongside-search and analysis summarisation. None of that is recoverable from the diff.

The spec carries the measurements and is explicit about which are too weak to decide on:

  • the answer-quality observation is three answers, no rubric — an anecdote
  • the determinism check is four inputs, and the failure mode it is meant to rule out is a rare flip, which is exactly what a small sample cannot rule out
  • price per token could not be read from any endpoint, so the premise that started this ("just as cheap") is still unverified

The blocker behind the blocker

src/evaluation/evaluator.py runs the right ragas metrics — faithfulness, answer_relevancy, context_recall, context_utilization — and cannot be used. It builds its own SelfQueryRetriever + EnsembleRetriever + MergerRetriever rather than calling the pipeline factory. Stage 2 removed SelfQueryRetriever from the product, so the evaluator now measures a configuration that no longer exists; running it would produce numbers that look like an answer and are not.

Same drift bin/retrieval_baseline had and had fixed. Fixing it is the spec P1 because it is owed regardless of which model wins: the four retrieval changes from 2026-09-04 went in unevaluated for answer quality.

Three decisions left to the team

Recorded rather than assumed — flip before or after evaluating (recommendation: after, and it is cheaper than it sounds), whether ~2x latency is acceptable on chat, and what luna actually costs.

Spec 001 closed out

Marked complete, with the outcome — and with the two predictions in its plan that turned out wrong:

  1. "Stage 1 must show zero difference" was unfalsifiable. Chroma ANN noise already differed on 2 of 8 question-collections between two identical runs, so zero could never have been observed — and a real regression that size would have been invisible against it.
  2. "Stage 2 becomes deterministic" did not happen. Removing the LLM from the vector side removed one of two sources of variance; the plan asserted the effect would vanish without checking which cause dominated.

Recorded because the point of these files is that the next person does not repeat them.

Two pieces of design tracking that were owed.

Spec 002 records the gpt-5.6-luna decision. #186 made the model usable;
this is about whether it becomes the default, which is not a dependency
bump: it changes what every user reads, trades determinism for
capability, doubles latency, moves cost in an unmeasured direction, and
will be asked again for chat-alongside-search and analysis
summarisation. None of that is recoverable from the diff.

It carries the measurements taken so far and is explicit about which of
them are too weak to decide on -- the answer-quality observation is
three answers with no rubric, and the determinism check is four inputs,
which cannot rule out the rare flip it is meant to rule out. Price per
token could not be read from any endpoint, so the premise that started
this ("just as cheap") is still unverified.

It also names the blocker behind the blocker: evaluator.py runs the
right ragas metrics but builds its own SelfQueryRetriever +
EnsembleRetriever + MergerRetriever rather than the shipping pipeline.
Stage 2 removed SelfQuery, so it now measures a configuration that does
not exist, and running it would produce numbers that look like an answer
and are not. Fixing that is the spec's P1 because it is owed anyway:
the four retrieval changes from 2026-09-04 went in unevaluated.

Three decisions are left to the team rather than assumed -- flip before
or after evaluating, whether 2x latency is acceptable on chat, and what
luna actually costs.

Spec 001 is marked complete with its outcome, including the two
predictions in its plan that were wrong: "Stage 1 must show zero
difference" was unfalsifiable against Chroma's ANN noise floor, which
already differed on 2 of 8 question-collections between identical runs;
and "Stage 2 becomes deterministic" did not happen, because removing the
LLM from the vector side removed only one of the two sources of
variance.
@adamjohnwright
adamjohnwright merged commit 6f9dc7c into main Sep 9, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the spec/default-llm-choice branch September 9, 2026 14:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant