Skip to content

Point the evaluator at the pipeline that ships - #191

Merged
adamjohnwright merged 1 commit into
mainfrom
fix/evaluator-measures-the-pipeline
Sep 9, 2026
Merged

adamjohnwright merged 1 commit into
mainfrom
fix/evaluator-measures-the-pipeline

Conversation

@adamjohnwright

Copy link
Copy Markdown
Contributor

Spec 002 P1 (FR-001..FR-004). Unblocks the gpt-5.6-luna decision and the four retrieval changes from 2026-09-04 that shipped unevaluated.

The problem

src/evaluation/evaluator.py ran the right ragas metrics and measured the wrong thing. It built its own retriever — SelfQueryRetriever + EnsembleRetriever + MergerRetriever, over the summations collection alone (of four), with k=7 and weights [0.2, 0.8] that appear nowhere in the product.

The retriever rewrite then removed SelfQueryRetriever from the pipeline, so it measured a configuration that existed nowhere at all. Running it would have produced numbers that looked like an answer.

It now calls create_reactome_rag — the same factory bin/chat-chainlit.py reaches. --rag_type basic|advanced is gone: those were two shapes of the private stack, not of the product. The embedding comes from resolve_embedding_model() rather than a hardcoded text-embedding-3-large, so Plant Reactome is not broken by it.

./bin/evaluate --model gpt-4o-mini --model gpt-5.6-luna --repeat 3

The judge is pinned separately (gpt-4o), and the tool refuses to run when the judge is also under test — a model grading its own answers is not a measurement. --repeat reports the spread as the noise floor, because retrieval is not deterministic and a single run cannot tell a real difference from Chroma ANN variance. The report says so rather than letting the reader assume otherwise.

Five things an adversarial review of my own rewrite caught

All fixed before this landed:

private API the aggregate came from result._repr_dict — in a file whose job is to stay trustworthy across upgrades. Now result.scores, which is public.
NaN poisoning a metric failing on one sample returns NaN, and NaN propagates through fmean — one bad question turned the whole aggregate into nan. Non-finite scores are dropped and the drop reported.
base_url hazard the judge embedding was a bare text-embedding-3-large with no base_url, so it followed OPENAI_BASE_URL. On the Plant Reactome host that points at a self-hosted bge-m3 endpoint — a 404 mid-run. Now api.openai.com explicitly, JUDGE_BASE_URL to override.
lost answers only scores were kept. The version this replaces wrote responses to a spreadsheet; dropping that leaves a low faithfulness score with nothing to look at. --out now carries every answer with its per-question scores.
stale README still documented --testset_dir and --rag_type. Rewritten, including why those flags are gone.

Verified by running it

Against the Release95 bundle, real scores from the real pipeline:

2 questions, judged by gpt-4o, 1 run(s) each

  model               answer_relevancy  context_utilizat      faithfulness   s/question
  gpt-4o-mini                    0.881             0.192             0.815         18.2

And the guard refusing --model gpt-4o --judge-model gpt-4o.

One thing the first runs already suggest

context_utilization came out 0.08–0.27 across the questions tried. If that holds over the full golden set, most of the 40 retrieved documents are contributing nothing — which speaks directly to the question spec 001 left open ("This value is a starting point, not a tuned one"). Too few questions to claim it yet; it is now measurable, which is the point.

src/evaluation/evaluator.py runs the right ragas metrics and measured
the wrong thing. It built its own retriever -- SelfQueryRetriever +
EnsembleRetriever + MergerRetriever, over the `summations` collection
alone, with k=7 and weights [0.2, 0.8] that appear nowhere in the
product. The retriever rewrite then removed SelfQueryRetriever from the
pipeline, so it measured a configuration that existed nowhere at all.
Running it would have produced numbers that looked like an answer.

It now calls create_reactome_rag, the same factory
bin/chat-chainlit.py reaches: four collections, the real budget, the
real fusion. There is one pipeline, so --rag_type basic|advanced is
gone -- those were two shapes of the private stack, not of the product.
The embedding comes from resolve_embedding_model() rather than a
hardcoded text-embedding-3-large, so Plant Reactome is not broken by it.

This is spec 002's P1 (FR-001..FR-004). It unblocks the gpt-5.6-luna
decision and the four retrieval changes from 2026-09-04 that shipped
unevaluated.

  ./bin/evaluate --model gpt-4o-mini --model gpt-5.6-luna --repeat 3

The model under test is an argument and can be repeated; the judge is
pinned separately, defaults to gpt-4o, and the tool refuses to run when
the judge is also under test, because a model grading its own answers
is not a measurement. --repeat reports the spread across runs as the
noise floor: retrieval is not deterministic, so a single run cannot
distinguish a real difference from Chroma's ANN variance, and the
report says so rather than letting the reader assume otherwise.

Five things an adversarial review of the rewrite caught, all fixed
before this landed:

- the aggregate came from result._repr_dict, a private ragas attribute,
  in a file whose job is to stay trustworthy across upgrades. Now
  result.scores, which is public.
- a metric that fails on one sample returns NaN, and NaN propagates
  through fmean -- one bad question would have turned the whole
  aggregate into "nan". Non-finite scores are dropped and the drop is
  reported.
- the judge's embedding was a bare "text-embedding-3-large" with no
  base_url, so it would follow OPENAI_BASE_URL. On the Plant Reactome
  host that points at a self-hosted bge-m3 endpoint: a 404 mid-run.
  api.openai.com is now named explicitly, with JUDGE_BASE_URL to
  override.
- only scores were kept, not answers. The version this replaces wrote
  responses to a spreadsheet; dropping that would have left a low
  faithfulness score with nothing to look at. --out now carries every
  answer with its per-question scores.
- the README still documented --testset_dir and --rag_type. Rewritten,
  including why those flags are gone.

Verified by running it: 2 and 3 question sets against the Release95
bundle, real scores, and the judge guard refusing gpt-4o vs gpt-4o.
@adamjohnwright
adamjohnwright merged commit 9d88a14 into main Sep 9, 2026
10 checks passed
@adamjohnwright
adamjohnwright deleted the fix/evaluator-measures-the-pipeline branch September 9, 2026 16:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant