Add nadir-caliper: consensus-verified cascade from Nadir Lab (true grok-4.3 prices; supersedes nadir-tumbler) - #173
Open
doramirdor wants to merge 1 commit into
Conversation
…des nadir-tumbler) Rerun at true grok-4.3 prices per the RouteWorks#166/RouteWorks#172 invitation. Two cheap probes (qwen3-235b + gemini-3.1-flash-lite), agreement gate, grok-4.3 arbiter. All fitting on RouterBench only, frozen before routing; per-row provenance (requested_model, model_used, provider, invoked_at) on new generations. Also registers grok-4.3 in universal_model_names.py, which RouteWorks#172 priced but did not register. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
Author
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Evaluation completed by RouterArena automated workflow |
Contributor
Author
Contributor
Author
|
Following up on our earlier comment: #174 evaluated below this submission (74.10 vs 74.55), so we have withdrawn it. Please disregard the supersede note above; this PR (nadir-caliper, 74.55) is Nadir Lab's entry and is ready to merge whenever convenient. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Submission from Nadir Lab (https://getnadir.com, https://github.com/NadirRouter, info@getnadir.com).
This is the rerun invited in #166 / #172 after the grok-4.1-fast retirement repricing. It supersedes our
nadir-tumblerentry; we would appreciate the leaderboard keeping a single Nadir Lab row pointed at this submission, with the affiliation shown as Nadir Lab linking to https://getnadir.com.What the router is
Menu:
qwen/qwen3-235b-a22b-2507,google/gemini-3.1-flash-lite,grok-4.3, plusdeepseek/deepseek-v3.2only as a declared per-row probe fallback (44 rows, see below).Decision rule per prompt (fully deterministic, no learned scores at inference):
Shipped distribution: qwen 74.6%, grok-4.3 16.0%, gemini-3.1-flash-lite 9.4%, deepseek-v3.2 4 rows. The
costfields sum to $1.814978 ($0.2161/1K) at the currentmodel_cost.jsonprices, computed from recorded token usage.Training data policy (evaluation-only compliance)
Provenance (adopting the guardrails proposed in #166)
requested_model,model_used,provider, andinvoked_atinsidegenerated_result, via OpenRouter, which echoes the served model.grok-4-1-fast-reasoningslug after xAI's May 15 redirect and therefore resolved to grok-4.3 with low reasoning effort, as established in Leaderboard integrity: retired Grok 4.1 Fast slug silently resolved to Grok 4.3 at the old price #166. They are declared asgrok-4.3and costed at grok-4.3 prices. We match the same low reasoning effort on all new grok-4.3 calls for consistency.Note for maintainers: grok-4.3 registration
PR #172 added grok-4.3 to
model_cost.jsonbutuniversal_model_names.pydoes not register it, so the submission preflight rejects the very model the corrected leaderboard prices. This PR includes the one-line registration (aftergrok-4-1-fast-reasoning). Cross-Router and vLLM-SR will need the same line for their invited reruns; feel free to take that change separately if you prefer.Robustness
nadir-caliper-robustness.jsonis a genuine re-run of the identical rule on the 420 perturbed prompts, with freshly generated probe responses (both probes) and freshly generated arbiter responses wherever the perturbed probes disagreed. Decision flips relative to the main split: 85 of 420 (flip-stability 79.76% byglobal_utils/robustness.py). No decisions were copied from the main split.Cost accounting disclosure
As with our previous submission, the benchmark charges only the shipped model. For transparency: in production this router pays for both probes on every query and the arbiter on the 30.6% of queries where the probes disagree, about $0.47 per 1K queries at list prices, versus $0.76 for always running grok-4.3. The benchmark-charged figure ($0.2161/1K) reflects the shipped models only.
On the accounting discrepancy flagged in #166 for our previous artifact: the committed file's
costfields sum to $0.887913, andcompute_scores.pyreproduces that number, while the PR evaluation recorded $0.633785. We believe the delta is the evaluator recomputing output tokens from visible text, which drops grok's billed reasoning tokens. This artifact states its exact sum above so there is no ambiguity; we are happy to help reconcile either way.Files
router_inference/config/nadir-caliper.jsonrouter_inference/predictions/nadir-caliper.json(8,400 rows,generated_resultpopulated, provenance fields on 2026-08-05 rows)router_inference/predictions/nadir-caliper-robustness.json(420 rows, decisions only)universal_model_names.py(one line: registergrok-4.3)Both
check_config_prediction_files.pygates pass locally (fullandrobustness).Contact: info@getnadir.com. Thanks again to the maintainers and the #166 auditors; the repricing was fair, and the provenance fields in this artifact are our attempt to make the next audit trivial.
🤖 Generated with Claude Code