Skip to content

Forward programme: synthesiser model/effort validation #65

Description

@Jodre11

Context

Thread 4 of the code-review-suite forward programme (effectiveness, agreed 2026-06-17).

Assumption under test

opus + ultrathink is the right call for the synthesiser. It's probably the most justified place for the highest tier (genuine judgement synthesis, cross-domain reasoning, verdict computation) — but "probably right" ≠ "empirically validated". The synthesiser is the single most expensive call in the pipeline; even a one-tier downgrade that holds quality compounds meaningfully.

Hard problem

How to score synthesis quality mechanically without a model-as-judge (explicitly rejected for the A/B harness). The per-specialist harness scores at a planted line; synthesis has no planted line. Needs a new scoring methodology — candidates:

  • Verdict stability — does a lower tier flip the verdict on the same specialist input?
  • Finding retention — does it drop/add findings the specialists produced?
  • Human preference ranking — small sample, but honest.

Why this ordering

After thread 3: the per-specialist sweep refines the measurement approach; apply those lessons to the harder synthesiser case.

Status

Not started. Hardest measurement problem in the programme; depends on thread 3's methodology.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions