Repository navigation
Spec 002: choosing the default answering model; close out spec 001 - #188
Merged
Merged
Conversation
Two pieces of design tracking that were owed. Spec 002 records the gpt-5.6-luna decision. #186 made the model usable; this is about whether it becomes the default, which is not a dependency bump: it changes what every user reads, trades determinism for capability, doubles latency, moves cost in an unmeasured direction, and will be asked again for chat-alongside-search and analysis summarisation. None of that is recoverable from the diff. It carries the measurements taken so far and is explicit about which of them are too weak to decide on -- the answer-quality observation is three answers with no rubric, and the determinism check is four inputs, which cannot rule out the rare flip it is meant to rule out. Price per token could not be read from any endpoint, so the premise that started this ("just as cheap") is still unverified. It also names the blocker behind the blocker: evaluator.py runs the right ragas metrics but builds its own SelfQueryRetriever + EnsembleRetriever + MergerRetriever rather than the shipping pipeline. Stage 2 removed SelfQuery, so it now measures a configuration that does not exist, and running it would produce numbers that look like an answer and are not. Fixing that is the spec's P1 because it is owed anyway: the four retrieval changes from 2026-09-04 went in unevaluated. Three decisions are left to the team rather than assumed -- flip before or after evaluating, whether 2x latency is acceptable on chat, and what luna actually costs. Spec 001 is marked complete with its outcome, including the two predictions in its plan that were wrong: "Stage 1 must show zero difference" was unfalsifiable against Chroma's ANN noise floor, which already differed on 2 of 8 question-collections between identical runs; and "Stage 2 becomes deterministic" did not happen, because removing the LLM from the vector side removed only one of the two sources of variance.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two pieces of design tracking that were owed.
Spec 002 — the gpt-5.6-luna decision
#186 made the model usable. This is about whether it becomes the default, which is not a dependency bump — it changes what every user reads, trades determinism for capability, doubles latency, moves cost in an unmeasured direction, and the same question returns for chat-alongside-search and analysis summarisation. None of that is recoverable from the diff.
The spec carries the measurements and is explicit about which are too weak to decide on:
The blocker behind the blocker
src/evaluation/evaluator.pyruns the right ragas metrics —faithfulness,answer_relevancy,context_recall,context_utilization— and cannot be used. It builds its ownSelfQueryRetriever+EnsembleRetriever+MergerRetrieverrather than calling the pipeline factory. Stage 2 removedSelfQueryRetrieverfrom the product, so the evaluator now measures a configuration that no longer exists; running it would produce numbers that look like an answer and are not.Same drift
bin/retrieval_baselinehad and had fixed. Fixing it is the spec P1 because it is owed regardless of which model wins: the four retrieval changes from 2026-09-04 went in unevaluated for answer quality.Three decisions left to the team
Recorded rather than assumed — flip before or after evaluating (recommendation: after, and it is cheaper than it sounds), whether ~2x latency is acceptable on chat, and what luna actually costs.
Spec 001 closed out
Marked complete, with the outcome — and with the two predictions in its plan that turned out wrong:
Recorded because the point of these files is that the next person does not repeat them.