fix: forward experimentIds from every model benchmark - #123
Conversation
|
I'll fix CI failures and address comments from users with write access that start with 'Devin'.
Original prompt from Abhinav
|
There was a problem hiding this comment.
Devin Review found 1 potential issue.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| providerIgnore: benchmarkConfig.providerIgnore, | ||
| allowFallbacks: benchmarkConfig.allowFallbacks, | ||
| cloudflareVersion: benchmarkConfig.cloudflareVersion, | ||
| experimentIds: benchmarkConfig.experimentIds, |
There was a problem hiding this comment.
🔴 CLI agent experiments use the default route
With an Ori agent, experimentIds reaches inference but not the agent's requests. buildAgentCliEnv omits them, so deep-swe and swe-atlas measure the default route.
Learn more
Deep-swe and swe-atlas support both direct Responses requests and external Ori agents. The added inference option reaches the direct model calls, but the Ori branch in makeDeepSweSolver runs runAgentCli instead. That runner builds the agent's environment through buildAgentCliEnv, which has no experiment setting. The same split applies to makeSweAtlasSolver; terminal-bench is also an Ori-only benchmark.
Example: Run deep-swe with agent: "ori" (or another valid Ori agent) and experimentIds: ["test-arm"]. The direct-model branch would send X-OpenRouter-Experiment-Ids: test-arm; the Ori agent's model requests receive no experiment setting.
Recommended fix: Pass experiment IDs through the CLI options and a supported agent configuration or request-header mechanism. Cover the Ori path in deep-swe and swe-atlas, and check terminal-bench's shared runner before claiming coverage for every model benchmark.
Was this helpful? React with 👍 or 👎 to provide feedback.
TL;DR
experimentIdsnow reaches the model request in every benchmark that makes direct Responses calls, not only tau-bench airline and tau3 banking.What changed?
experimentIdsnext tocloudflareVersion. The Responses model already turns it intoX-OpenRouter-Experiment-Ids.searchSolverOptionsFromConfignow forwardsexperimentIds, andsearchSolversends it asX-OpenRouter-Experiment-Ids. Judge and grader requests are unchanged.Why?
Follow-up to #122, which Devin Review flagged on https://github.com/OpenRouterTeam/openrouter-web/pull/47401. The shared
InferenceOverrideSchemaacceptedexperimentIdsfor all benchmarks, but most builders dropped it, so those runs silently measured the default route.How to test
solver.test.tscase checks that search requests carryX-OpenRouter-Experiment-Ids: jev-finish-deesc,control.bun src/cli/index.ts --benchmark gpqa_diamond --model typesafe/jev-router --limit 1 --solver-config '{"experimentIds":["jev-finish-deesc"]}': the generation should resolve to the arm's pool.Reviewer focus
agentset, and terminal-bench).buildAgentCliEnvhas no experiment setting, so those agents still use the default route. Fixing that needs an agent-side header mechanism.Checklist
Link to Devin session: https://openrouter.devinenterprise.com/sessions/6b3d4738964a49649c5e97ec44a18f3e
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/6b3d4738964a49649c5e97ec44a18f3e?variant=devin
Requested by: @abhinav-pola