Context
Thread 3 of the code-review-suite forward programme (effectiveness, agreed 2026-06-17).
Goal
Find the optimum (model, effort) for each agentic/judgement specialist independently. Unlike the five static specialists (all haiku/low — they wrap deterministic tools), the agentic specialists do open-ended reasoning and may each land at a different answer.
In scope (7): correctness, security, style, reuse, alignment, archaeology, test-quality.
Apparatus: the existing per-agent A/B harness + corpus fixtures.
Hypotheses to test (empirical, not assumed)
- Fable for security — likely best for adversarial reasoning, BUT the read-only guardrails wrapped around specialists might hold it back. Test, don't assume. (Gated on Fable availability; the rest of the sweep is testable now.)
- Haiku for alignment/archaeology — possibly simpler judgement tasks that don't need frontier reasoning.
- Sonnet as baseline — all agentic specialists currently run sonnet/default. Is that right, or over/under-specified for some?
Why this ordering
After thread 2 (phase-efficacy): if cross-review discards most of a specialist's output, optimising that specialist's model is lower-leverage than fixing pipeline structure.
Status
Not started. Depends on thread 2 findings + the A/B harness.
Context
Thread 3 of the code-review-suite forward programme (effectiveness, agreed 2026-06-17).
Goal
Find the optimum
(model, effort)for each agentic/judgement specialist independently. Unlike the five static specialists (all haiku/low — they wrap deterministic tools), the agentic specialists do open-ended reasoning and may each land at a different answer.In scope (7): correctness, security, style, reuse, alignment, archaeology, test-quality.
Apparatus: the existing per-agent A/B harness + corpus fixtures.
Hypotheses to test (empirical, not assumed)
Why this ordering
After thread 2 (phase-efficacy): if cross-review discards most of a specialist's output, optimising that specialist's model is lower-leverage than fixing pipeline structure.
Status
Not started. Depends on thread 2 findings + the A/B harness.