Skip to content

fix: forward experimentIds from every model benchmark - #123

Merged
abhinav-pola merged 1 commit into
mainfrom
devin/1790484707-experiment-ids-all-benchmarks
Sep 27, 2026
Merged

abhinav-pola merged 1 commit into
mainfrom
devin/1790484707-experiment-ids-all-benchmarks

Conversation

@abhinav-pola

@abhinav-pola abhinav-pola commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

TL;DR

experimentIds now reaches the model request in every benchmark that makes direct Responses calls, not only tau-bench airline and tau3 banking.

What changed?

  • The inference builders for gpqa, mmlu-pro, mmmu-pro-vision, ifstruct, wandr, vgi-bench, deep-swe and swe-atlas now copy experimentIds next to cloudflareVersion. The Responses model already turns it into X-OpenRouter-Experiment-Ids.
  • Search benchmarks (hle, dsqa, widesearch, browsecomp): searchSolverOptionsFromConfig now forwards experimentIds, and searchSolver sends it as X-OpenRouter-Experiment-Ids. Judge and grader requests are unchanged.

Why?

Follow-up to #122, which Devin Review flagged on https://github.com/OpenRouterTeam/openrouter-web/pull/47401. The shared InferenceOverrideSchema accepted experimentIds for all benchmarks, but most builders dropped it, so those runs silently measured the default route.

How to test

  • The new solver.test.ts case checks that search requests carry X-OpenRouter-Experiment-Ids: jev-finish-deesc,control.
  • bun src/cli/index.ts --benchmark gpqa_diamond --model typesafe/jev-router --limit 1 --solver-config '{"experimentIds":["jev-finish-deesc"]}': the generation should resolve to the arm's pool.

Reviewer focus

  • Grader and judge calls deliberately don't get the header.
  • Not covered: Ori/agent-CLI runs (deep-swe or swe-atlas with agent set, and terminal-bench). buildAgentCliEnv has no experiment setting, so those agents still use the default route. Fixing that needs an agent-side header mechanism.

Checklist

  • Tests cover changed behavior
  • Public API or configuration changes are backward compatible, or the break is documented
  • No credentials, private results, or restricted dataset contents are included

Link to Devin session: https://openrouter.devinenterprise.com/sessions/6b3d4738964a49649c5e97ec44a18f3e
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/6b3d4738964a49649c5e97ec44a18f3e?variant=devin
Requested by: @abhinav-pola


Devin Review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

Original prompt from Abhinav

SYSTEM:
<latest_message>
Abhinav Pola (U090K0G7JF3) [ts=1790445587.425579]: @Devin
DO NOT LOOK AT ANY RESULTS IN THE BENCHMARK RESULTS TABLE
looking at the deepswe tasks (without looking at any existing results), can you guess the difficulty of each task? produce a TSV

  • Read the hidden tests
  • Check which requirements the instructions leave out
  • Look at the actual repo, not just the diff
  • Easiest: narwhals-rolling-window-suite. 7/7 on the strong-model runs and 18/18 across all 23 full runs. It's the only task every graded run solved.
  • Hardest: igel-persist-feature-schema. 0/7 on the strong-model runs and 0/26 across all runs.
    </latest_message>

=== BEGIN THREAD HISTORY (in #agents-benchmarks) ===
Abhinav Pola (U090K0G7JF3) [ts=1790445587.425579]: @Devin
DO NOT LOOK AT ANY RESULTS IN THE BENCHMARK RESULTS TABLE
looking at the deepswe tasks (without looking at any existing results), can you guess the difficulty of each task? produce a TSV

The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

providerIgnore: benchmarkConfig.providerIgnore,
allowFallbacks: benchmarkConfig.allowFallbacks,
cloudflareVersion: benchmarkConfig.cloudflareVersion,
experimentIds: benchmarkConfig.experimentIds,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 CLI agent experiments use the default route

With an Ori agent, experimentIds reaches inference but not the agent's requests. buildAgentCliEnv omits them, so deep-swe and swe-atlas measure the default route.

Learn more

Deep-swe and swe-atlas support both direct Responses requests and external Ori agents. The added inference option reaches the direct model calls, but the Ori branch in makeDeepSweSolver runs runAgentCli instead. That runner builds the agent's environment through buildAgentCliEnv, which has no experiment setting. The same split applies to makeSweAtlasSolver; terminal-bench is also an Ori-only benchmark.

Example: Run deep-swe with agent: "ori" (or another valid Ori agent) and experimentIds: ["test-arm"]. The direct-model branch would send X-OpenRouter-Experiment-Ids: test-arm; the Ori agent's model requests receive no experiment setting.

Recommended fix: Pass experiment IDs through the CLI options and a supported agent configuration or request-header mechanism. Cover the Ori path in deep-swe and swe-atlas, and check terminal-bench's shared runner before claiming coverage for every model benchmark.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@abhinav-pola
abhinav-pola merged commit 232356d into main Sep 27, 2026
5 checks passed
@abhinav-pola
abhinav-pola deleted the devin/1790484707-experiment-ids-all-benchmarks branch September 27, 2026 04:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant