Skip to content

ci(qp-topic-label): offer claude-opus-5-5; claude CLI pin → 2.1.280 - #1962

Merged
modelmirror merged 1 commit into
stagingfrom
feat/qp-label-opus-5-5
Sep 22, 2026
Merged

modelmirror merged 1 commit into
stagingfrom
feat/qp-label-opus-5-5

Conversation

@modelmirror

Copy link
Copy Markdown
Collaborator

What changed and why

Adds claude-opus-5-5 (Claude Opus 5.5) as a label_model choice for the qp-topic-label mode of run-analytics. The default stays Haiku.

Opus 5.5 needs Claude Code 2.1.280 or newer. Older CLIs reject the model id with a 400 ("version 2.1.280 or newer is required"). The labeler installs its own pinned CLI (@anthropic-ai/claude-code@2.1.259), so the pin moves to 2.1.280. The comment in run-analytics says the pin is kept in version step with run-backtest and integration-test, and a test enforces a single pin across the three workflows. So the change touches all four places:

  • run-analytics.yml: the labeler install
  • integration-test.yml: the engine-smoke CLI install and qp-labeler-smoke
  • run-backtest.yml: the replay CLI install

The claude-code-action pin (v1.0.217) is deliberately not bumped here. The labeler passes it path_to_claude_code_executable, so the CLI bundled with the action goes unused. Production predict/evaluate cells keep running on that bundled CLI. Moving them to Opus 5.5 is a separate change, staged in its own PR, which waits until after the long conference.

Artifact served: the qp-topic labels artifact (data/qp-topics/). This PR adds a model option; it produces nothing on its own.

Comments updated:

  • The run-analytics comment about the action "bundling a newer CLI" was now backwards. The pin now runs ahead of the bundled CLI (about 2.1.263).
  • The integration-test lockstep comment named the wrong workflows for the claude pin.
  • docs/pipeline.md gains one clause: Opus 5.5 is offered but has no measured pace or cost yet.

Risk, and the checks that settle it (maintainer dispatches)

The action's agent SDK (0.3.263) will now drive a 2.1.280 CLI. A newer executable is usually compatible, but the SDK↔CLI protocol is not a documented stable contract. The labeler's count-keyed publication and agreement gate keep bad labels from landing, but a mismatch could still waste a run. This is unverified until these run:

  1. After merge to staging, run the labeler smoke. It exercises the exact action + 2.1.280 CLI pairing, on the Haiku default:
    gh workflow run integration-test.yml --repo ModelMirrorAI/fedcourtsai --ref staging -f scenario=qp-labeler-smoke
  2. Rehearse Opus 5.5 itself on staging. Nothing else smokes this model. The run publishes nothing to prod:
    gh workflow run run-analytics.yml --repo ModelMirrorAI/fedcourtsai --ref staging -f mode=qp-topic-label -f label_model=claude-opus-5-5
    Expected: the resolve/install step prints 2.1.280 (Claude Code), the agent step runs on claude-opus-5-5, and the measure step reports an agreement figure.

The CLI bump also reaches the backtest's claude replay cells and the claude engine-smoke leg. engine-smoke on staging covers that pairing.

Promotion effect check: on a prod qp-topic-label dispatch with label_model=claude-opus-5-5, the labels artifact's batch records labeler: claude-code/claude-opus-5-5.

Review

workflow-reviewer verdict: recommended, no blockers. Its two comment findings are fixed. The SDK/CLI skew risk is the dispatch plan above. zizmor 1.26.1 and actionlint are clean, and tests/test_workflow_cell_invariants.py passes (76).

This PR is under .github/workflows/, so it waits for the maintainer; I have not merged it.

🤖 Generated with Claude Code

… to 2.1.280

Opus 5.5 needs Claude Code 2.1.280 or newer; the pin moves in step across
run-analytics, run-backtest and integration-test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant