Point the evaluator at the pipeline that ships - #191
Merged
Merged
Conversation
src/evaluation/evaluator.py runs the right ragas metrics and measured the wrong thing. It built its own retriever -- SelfQueryRetriever + EnsembleRetriever + MergerRetriever, over the `summations` collection alone, with k=7 and weights [0.2, 0.8] that appear nowhere in the product. The retriever rewrite then removed SelfQueryRetriever from the pipeline, so it measured a configuration that existed nowhere at all. Running it would have produced numbers that looked like an answer. It now calls create_reactome_rag, the same factory bin/chat-chainlit.py reaches: four collections, the real budget, the real fusion. There is one pipeline, so --rag_type basic|advanced is gone -- those were two shapes of the private stack, not of the product. The embedding comes from resolve_embedding_model() rather than a hardcoded text-embedding-3-large, so Plant Reactome is not broken by it. This is spec 002's P1 (FR-001..FR-004). It unblocks the gpt-5.6-luna decision and the four retrieval changes from 2026-09-04 that shipped unevaluated. ./bin/evaluate --model gpt-4o-mini --model gpt-5.6-luna --repeat 3 The model under test is an argument and can be repeated; the judge is pinned separately, defaults to gpt-4o, and the tool refuses to run when the judge is also under test, because a model grading its own answers is not a measurement. --repeat reports the spread across runs as the noise floor: retrieval is not deterministic, so a single run cannot distinguish a real difference from Chroma's ANN variance, and the report says so rather than letting the reader assume otherwise. Five things an adversarial review of the rewrite caught, all fixed before this landed: - the aggregate came from result._repr_dict, a private ragas attribute, in a file whose job is to stay trustworthy across upgrades. Now result.scores, which is public. - a metric that fails on one sample returns NaN, and NaN propagates through fmean -- one bad question would have turned the whole aggregate into "nan". Non-finite scores are dropped and the drop is reported. - the judge's embedding was a bare "text-embedding-3-large" with no base_url, so it would follow OPENAI_BASE_URL. On the Plant Reactome host that points at a self-hosted bge-m3 endpoint: a 404 mid-run. api.openai.com is now named explicitly, with JUDGE_BASE_URL to override. - only scores were kept, not answers. The version this replaces wrote responses to a spreadsheet; dropping that would have left a low faithfulness score with nothing to look at. --out now carries every answer with its per-question scores. - the README still documented --testset_dir and --rag_type. Rewritten, including why those flags are gone. Verified by running it: 2 and 3 question sets against the Release95 bundle, real scores, and the judge guard refusing gpt-4o vs gpt-4o.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Spec 002 P1 (FR-001..FR-004). Unblocks the
gpt-5.6-lunadecision and the four retrieval changes from 2026-09-04 that shipped unevaluated.The problem
src/evaluation/evaluator.pyran the right ragas metrics and measured the wrong thing. It built its own retriever —SelfQueryRetriever+EnsembleRetriever+MergerRetriever, over thesummationscollection alone (of four), withk=7and weights[0.2, 0.8]that appear nowhere in the product.The retriever rewrite then removed
SelfQueryRetrieverfrom the pipeline, so it measured a configuration that existed nowhere at all. Running it would have produced numbers that looked like an answer.It now calls
create_reactome_rag— the same factorybin/chat-chainlit.pyreaches.--rag_type basic|advancedis gone: those were two shapes of the private stack, not of the product. The embedding comes fromresolve_embedding_model()rather than a hardcodedtext-embedding-3-large, so Plant Reactome is not broken by it.The judge is pinned separately (
gpt-4o), and the tool refuses to run when the judge is also under test — a model grading its own answers is not a measurement.--repeatreports the spread as the noise floor, because retrieval is not deterministic and a single run cannot tell a real difference from Chroma ANN variance. The report says so rather than letting the reader assume otherwise.Five things an adversarial review of my own rewrite caught
All fixed before this landed:
result._repr_dict— in a file whose job is to stay trustworthy across upgrades. Nowresult.scores, which is public.fmean— one bad question turned the whole aggregate intonan. Non-finite scores are dropped and the drop reported.text-embedding-3-largewith nobase_url, so it followedOPENAI_BASE_URL. On the Plant Reactome host that points at a self-hosted bge-m3 endpoint — a 404 mid-run. Nowapi.openai.comexplicitly,JUDGE_BASE_URLto override.--outnow carries every answer with its per-question scores.--testset_dirand--rag_type. Rewritten, including why those flags are gone.Verified by running it
Against the Release95 bundle, real scores from the real pipeline:
And the guard refusing
--model gpt-4o --judge-model gpt-4o.One thing the first runs already suggest
context_utilizationcame out 0.08–0.27 across the questions tried. If that holds over the full golden set, most of the 40 retrieved documents are contributing nothing — which speaks directly to the question spec 001 left open ("This value is a starting point, not a tuned one"). Too few questions to claim it yet; it is now measurable, which is the point.