Skip to content

Add nadir-caliper-2: consult-only arbiter (supersedes pending nadir-caliper #173) - #174

Closed
doramirdor wants to merge 1 commit into
RouteWorks:mainfrom
doramirdor:nadir-caliper-2-submission
Closed

Add nadir-caliper-2: consult-only arbiter (supersedes pending nadir-caliper #173)#174
doramirdor wants to merge 1 commit into
RouteWorks:mainfrom
doramirdor:nadir-caliper-2-submission

Conversation

@doramirdor

Copy link
Copy Markdown
Contributor

Submission from Nadir Lab (https://getnadir.com, https://github.com/NadirRouter, info@getnadir.com).

This supersedes our pending nadir-caliper submission (#173) and, like it, the nadir-tumbler leaderboard entry. One Nadir Lab row, please, with affiliation shown as Nadir Lab linking to https://getnadir.com. We are intentionally replacing a submission whose evaluation we have already seen (74.55) with a variant we have not: our RouterBench-only branch analysis, documented in the config description and committed before this PR, says this version dominates it. We accept the evaluated outcome either way.

The one change from nadir-caliper

Identical menu, probes, arbitration, and provenance. The only difference: where grok-4.3 confirms neither probe (the hardest 16% of traffic), the router ships qwen instead of grok. On our RouterBench labels that branch selects questions every model fails (grok correct 12%, qwen 15%), so the premium answer bought no accuracy at 25x the price. grok-4.3 is consult-only in this version and is never shipped; the >=30% downroute constraint is trivially satisfied at 100%.

Cost fields sum to $0.418165 ($0.0498/1K) at current model_cost.json prices from recorded token usage. The robustness file is a genuine re-run of the identical rule on all 420 perturbed prompts. Everything else (training-data policy, contamination audit, per-row provenance on 2026-08-05 generations, the grok-4.3 registry line) is unchanged from #173.

Files

Both check_config_prediction_files.py gates pass locally. Contact: info@getnadir.com.

🤖 Generated with Claude Code

… nadir-caliper, PR RouteWorks#173)

One branch changed from nadir-caliper: where grok-4.3 confirms neither probe,
ship qwen instead of grok. RouterBench-only evidence: that branch selects
questions all models fail (grok 12% vs qwen 15% correct), so the premium
answer bought no accuracy at 25x the price. grok-4.3 becomes consult-only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@doramirdor

Copy link
Copy Markdown
Contributor Author

/evaluate

@github-actions

github-actions Bot commented Aug 6, 2026

Copy link
Copy Markdown

Router Evaluation Results

Router: nadir-caliper-2
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7410
Accuracy 73.79%
Total Cost $0.418165
Avg Cost per Query $0.000050
Avg Cost per 1K Queries $0.0498
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8738

Evaluation completed by RouterArena automated workflow

@doramirdor

Copy link
Copy Markdown
Contributor Author

The evaluation answered the question this PR existed to ask: nadir-caliper-2 scores 74.10 vs nadir-caliper's 74.55, so the branch change is a measured regression on this dataset (the arbiter's answers on the no-confirmation slice are worth ~13pp there, unlike on RouterBench where we measured them at ~0). Per the decision rule stated in the PR body, we accept the evaluated outcome: withdrawing this submission and closing the PR. #173 (nadir-caliper) remains Nadir Lab's intended entry. Both artifacts and both evaluations stay public for anyone auditing the comparison. Thanks for the compute, and apologies again for the churn.

@doramirdor doramirdor closed this Aug 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant