Context
Thread 4 of the code-review-suite forward programme (effectiveness, agreed 2026-06-17).
Assumption under test
opus + ultrathink is the right call for the synthesiser. It's probably the most justified place for the highest tier (genuine judgement synthesis, cross-domain reasoning, verdict computation) — but "probably right" ≠ "empirically validated". The synthesiser is the single most expensive call in the pipeline; even a one-tier downgrade that holds quality compounds meaningfully.
Hard problem
How to score synthesis quality mechanically without a model-as-judge (explicitly rejected for the A/B harness). The per-specialist harness scores at a planted line; synthesis has no planted line. Needs a new scoring methodology — candidates:
- Verdict stability — does a lower tier flip the verdict on the same specialist input?
- Finding retention — does it drop/add findings the specialists produced?
- Human preference ranking — small sample, but honest.
Why this ordering
After thread 3: the per-specialist sweep refines the measurement approach; apply those lessons to the harder synthesiser case.
Status
Not started. Hardest measurement problem in the programme; depends on thread 3's methodology.
Context
Thread 4 of the code-review-suite forward programme (effectiveness, agreed 2026-06-17).
Assumption under test
opus + ultrathink is the right call for the synthesiser. It's probably the most justified place for the highest tier (genuine judgement synthesis, cross-domain reasoning, verdict computation) — but "probably right" ≠ "empirically validated". The synthesiser is the single most expensive call in the pipeline; even a one-tier downgrade that holds quality compounds meaningfully.
Hard problem
How to score synthesis quality mechanically without a model-as-judge (explicitly rejected for the A/B harness). The per-specialist harness scores at a planted line; synthesis has no planted line. Needs a new scoring methodology — candidates:
Why this ordering
After thread 3: the per-specialist sweep refines the measurement approach; apply those lessons to the harder synthesiser case.
Status
Not started. Hardest measurement problem in the programme; depends on thread 3's methodology.