Skip to content

Run each lens per model variant, scale to a granted time limit - #5

Draft
joahg wants to merge 2 commits into
mainfrom
six-panelist-variants
Draft

joahg wants to merge 2 commits into
mainfrom
six-panelist-variants

Conversation

@joahg

@joahg joahg commented Oct 10, 2026 •

Copy link
Copy Markdown
Member

Lets one review run each lens with more than one model, and with different providers, and makes gpt-6.1-sol usable through OpenAI's API.

Six-panelist panels. Config.PanelVariants runs every panelist lens once per named variant. With sol,opus, the review runs six panelists, for example security_contracts.sol and security_contracts.opus. Each panelist reports a name with its variant, such as "Security & Trust Boundaries (opus)". The coordinator prompt now lists the actual roster instead of hard-coding three panelists, and when variants are set it expects panelists that share a lens to report overlapping candidates. Without variants, the panel is the same three lenses as before.

Per-pass providers. Config.RoleProviders joins RoleModels and RoleEfforts, so the coordinator, the judge, and each panelist can use a different Goose provider. All three settings resolve from most to least specific: the pass role, then its variant, then its lens. For example, opus=anthropic covers all three Opus panelists, and security_contracts.opus=… overrides just one of them. The CLI adds --panel-variants and --role-providers, and the ReviewBench adapter adds RB_CONFIG_PANEL_VARIANTS and RB_CONFIG_ROLE_PROVIDERS.

Responses API for OpenAI. Goose 1.43 sends models it doesn't recognize, gpt-6.1-sol among them, to Chat Completions, which doesn't support tool calls for those models. When any pass uses the openai provider, the ReviewBench adapter sets OPENAI_BASE_PATH=v1/responses unless the caller set one. Against a local server, Goose then sent gpt-6.1-sol with its tools to /v1/responses instead of /v1/chat/completions.

Time limit. RB_CONFIG_TIME_LIMIT, in seconds and at least 900, names the per-PR limit granted in the ReviewBench manifest. Every stage scales from it. Without it, the budget is the 15-minute default as before.

Shorter coordinator wrap-up. A coordinator reconciling six panelists sometimes ran out of time, then ran out again while writing its long drop list in the wrap-up turn, which failed the pull request. On a time-up wrap-up, the coordinator now returns an empty dropped_candidates array.

Verification.

  • A new test runs a six-panelist review with mixed providers. It checks that every lens runs once per variant, that overrides resolve in specificity order, and that the coordinator prompt lists all six panelists.
  • go vet and go test pass. The default three-panelist coordinator prompt renders the same roster as before.
  • Running review reviewbench end to end against a fake goose produced six panelists: three on openai/gpt-6.1-sol and three on anthropic/claude-opus-5-5. Every call had the Responses base path set.
  • A new subtest checks that a coordinator's time-up wrap-up asks for an empty drop list. RB_CONFIG_TIME_LIMIT was checked by hand: unset gives 15m, 1800 gives 30m, and 600 or abc are rejected.

joahg added 2 commits October 9, 2026 22:50
PanelVariants runs every panelist lens once per variant, so variants sol
and opus give six panelists. RoleProviders joins RoleModels and
RoleEfforts, and all three resolve the most specific key first: the pass
role, then its variant, then its lens. The coordinator prompt now names
the actual roster and expects same-lens overlap when variants are set.

The ReviewBench adapter reads RB_CONFIG_PANEL_VARIANTS and
RB_CONFIG_ROLE_PROVIDERS, and sends OpenAI passes to the Responses API:
Goose routes models it does not recognize, such as gpt-6.1-sol, to Chat
Completions, where they cannot call tools.
…oordinator wrap-up

RB_CONFIG_TIME_LIMIT, in seconds, names the per-PR limit granted in the
manifest. Every stage scales with it; without it the budget is the
15-minute default as before.

A coordinator reconciling six panelists often ran out of time and then
ran out again writing its drop list in the wrap-up turn, failing the pull
request. On a time-up wrap-up the coordinator now returns an empty
dropped_candidates array.
@joahg joahg changed the title Run each lens per model variant, with per-pass providers Run each lens per model variant, scale to a granted time limit Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant