Skip to content

feat: send x-openrouter-experiment-ids from experimentIds option - #122

Merged
abhinav-pola merged 1 commit into
mainfrom
devin/1790483747-experiment-ids-header
Sep 27, 2026
Merged

abhinav-pola merged 1 commit into
mainfrom
devin/1790483747-experiment-ids-header

Conversation

@abhinav-pola

@abhinav-pola abhinav-pola commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

TL;DR

Adds an experimentIds inference option that the harness sends as x-openrouter-experiment-ids, so benchmark runs can opt into live-config jev_router experiment arms.

What changed?

  • InferenceOverrideSchema.experimentIds?: string[] (1–16 IDs, same pattern as the router's EXPERIMENT_NAME_PATTERN).
  • GenerateConfig.experimentIds → Responses request header X-OpenRouter-Experiment-Ids: id1,id2.
  • Wired through tau_bench_verified_airline and tau3_bench_banking agent inference. Other benchmarks can be wired the same way when needed.

Why?

The router only applies an experiment arm when the request sends this header (OpenRouter org keys only). Before this change there was no way to run a benchmark against an arm. First use: A/B testing jev-finish-deesc against jev-finish-deesc-control on tau-bench airline.

How to test

bun test src/providers/responses-model.test.ts
bun src/cli/index.ts --benchmark tau_bench_verified_airline --model typesafe/jev-router \
  --solver-config '{"experimentIds":["jev-finish-deesc"]}'

Reviewer focus

  • The header is omitted when experimentIds is unset, so default runs are unchanged.

Checklist

  • Tests cover changed behavior
  • Public API or configuration changes are backward compatible, or the break is documented
  • Benchmark changes document dataset provenance and licensing
  • No credentials, private results, or restricted dataset contents are included
  • Documentation is updated where needed

Link to Devin session: https://openrouter.devinenterprise.com/sessions/6b3d4738964a49649c5e97ec44a18f3e
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/6b3d4738964a49649c5e97ec44a18f3e?variant=devin
Requested by: @abhinav-pola


Devin Review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

Original prompt from Abhinav

SYSTEM:
<latest_message>
Abhinav Pola (U090K0G7JF3) [ts=1790445587.425579]: @Devin
DO NOT LOOK AT ANY RESULTS IN THE BENCHMARK RESULTS TABLE
looking at the deepswe tasks (without looking at any existing results), can you guess the difficulty of each task? produce a TSV

  • Read the hidden tests
  • Check which requirements the instructions leave out
  • Look at the actual repo, not just the diff
  • Easiest: narwhals-rolling-window-suite. 7/7 on the strong-model runs and 18/18 across all 23 full runs. It's the only task every graded run solved.
  • Hardest: igel-persist-feature-schema. 0/7 on the strong-model runs and 0/26 across all runs.
    </latest_message>

=== BEGIN THREAD HISTORY (in #agents-benchmarks) ===
Abhinav Pola (U090K0G7JF3) [ts=1790445587.425579]: @Devin
DO NOT LOOK AT ANY RESULTS IN THE BENCHMARK RESULTS TABLE
looking at the deepswe tasks (without looking at any existing results), can you guess the difficulty of each task? produce a TSV

The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

Comment on lines +58 to +62
experimentIds: z
.array(z.string().regex(/^[a-z0-9][a-z0-9_.-]{0,63}$/))
.min(1)
.max(16)
.optional(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Experiments silently ignored by other benchmarks

When other benchmarks set experimentIds, their configs validate but requests omit the experiment header. Their inference builders discard experimentIds, so runs measure the default arm instead.

Learn more

The shared inference schema feeds all model benchmark configurations, so parsing accepts experimentIds for benchmarks besides airline and banking. Most benchmark layers construct inference objects by explicitly copying individual fields, such as makeSolver and wandrInferenceOverride, and never copy the new field. Search benchmarks use searchSolverOptionsFromConfig and their own request path, which likewise omits the header. Validated runs therefore complete without selecting their requested experiment.

Example: A gpqa_diamond run with experimentIds: ["jev-finish-deesc"] passes validation, but its model request lacks X-OpenRouter-Experiment-Ids and evaluates the unselected arm.

Recommended fix: Either forward experimentIds through every benchmark that accepts the shared schema, including the search request path, or expose it only in the benchmark schemas whose inference paths support it. Add regression coverage from benchmark config through sent headers for the supported paths.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@abhinav-pola
abhinav-pola merged commit 84bcd9f into main Sep 27, 2026
5 checks passed
@abhinav-pola
abhinav-pola deleted the devin/1790483747-experiment-ids-header branch September 27, 2026 04:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants