Benchmark
150 real prompts, every one answered through the router and, for the baseline, by the frontier model; two frontier judges scored each pair blind, 1 to 10. Real provider calls at list price, 2026-09-10, OpenSmartRoute 1.1.0. The files behind every number are in the repository and the commands to re-run them are at the end of this page.
89.3%
less spent
$0.0263 routed against $0.2457 always-frontier, for 150 prompts
9.53 vs 9.15
judged quality, routed vs baseline
Mean of two blind judges, 1 to 10; routed answers won 68, tied 52 and lost 30
1.17 s vs 2.01 s
median latency, routed vs baseline
A smaller model answers most prompts faster
94%
of routed answers within one point of the baseline
The two judges agreed on the winner on 72% of prompts
Every model, directly, on the same prompts
Each row sent all 150 prompts to one model directly. The router row is the routed run - 122 prompts to a small model, 28 to a mid-size one - judged the same way. Sorted by quality.
| Option | Cost, 150 prompts | Per 1M requests | Quality (1-10) | p50 |
|---|---|---|---|---|
| Direct: gpt-oss-120b (AWS) | $0.0467 | $311.00 | 9.66 | 1.42 s |
| OpenSmartRoute, routedrouter | $0.0263 | $176.00 | 9.53 | 1.17 s |
| Direct: gpt-oss-20b (AWS) | $0.0168 | $112.00 | 9.52 | 1.33 s |
| Direct: gemini-2.5-pro (GCP) | $0.7696 | $5130.00 | 9.52 | 5.35 s |
| Direct: gpt-4.1-mini (Azure) | $0.0440 | $294.00 | 9.39 | 2.36 s |
| Direct: gpt-4.1 (Azure) - the reference | $0.2457 | $1638.00 | 9.27 | 2.01 s |
| Direct: gpt-4.1-nano (Azure) | $0.0109 | $73.00 | 9.17 | 2.18 s |
| Direct: gemini-2.5-flash (GCP) | $0.1282 | $855.00 | 9.11 | 1.86 s |
| Direct: gemini-2.5-flash-lite (GCP) | $0.0226 | $151.00 | 9.11 | 1.42 s |
Reference and judges: azure-gpt-4.1 answered every prompt as the baseline; azure-gpt-4.1 and gcp-gemini-2.5-pro judged each pair blind. One direct model scored higher than the router (9.66 for $0.0467); the router reached 9.53 for $0.0263 - on this catalogue it chose between exactly those two models, request by request. Routing changes which model bills the tokens, not how many are sent; reasoning models bill their thinking as output, so a cheaper row can show more tokens.
By use case
The same 150 prompts grouped by what they asked for. Savings are against the always-frontier baseline on the same prompts.
| Use case | Prompts | Routed quality | Baseline quality | Saved |
|---|---|---|---|---|
| Analysis | 11 | 9.82 | 8.77 | 90.3% |
| Chat | 15 | 9.30 | 8.80 | 86.6% |
| Classify | 13 | 9.19 | 9.88 | 93.3% |
| Code | 17 | 9.79 | 9.03 | 88.8% |
| Extract | 13 | 9.73 | 9.54 | 90.7% |
| FAQ | 14 | 9.43 | 9.54 | 89.2% |
| Legal and finance | 8 | 9.75 | 8.81 | 88.2% |
| Marketing | 11 | 9.50 | 8.77 | 85.0% |
| Math | 11 | 10.00 | 9.91 | 94.2% |
| Summarise | 13 | 9.54 | 9.15 | 86.2% |
| Support | 14 | 9.04 | 8.18 | 86.0% |
| Translate | 10 | 9.40 | 9.50 | 90.8% |
Classification and FAQ answers judged lower when routed to the small model - the cases a quality rule or a reported outcome would move to the larger one. The run was a cold start: no learned state, no feedback. The router learns from outcomes, so these are its worst numbers, not its best.
RouterBench, the hard one
RouterBench 0-shot: 18,239 held-out prompts, 11 real models with per-response correctness and cost. The catalogue and the routing model were built from the other half of the data; nothing below saw the prompts it is scored on. A selection of the policies - the full table and the reading are on the results page.
| Policy | Accuracy | Realised quality | $ per 1k calls |
|---|---|---|---|
| Oracle: the best model per prompt | 1.000 | 0.912 | $0.22 |
| Static task table, fitted on the train half | 0.830 | 0.780 | $2.98 |
| Always the most expensive model (gpt-4-1106-preview) | 0.819 | 0.780 | $3.31 |
| OpenSmartRoute, quality only (cost weight 0) | 0.810 | 0.786 | $3.25 |
| OpenSmartRoute, cost weight 0.1 | 0.805 | - | $2.08 |
| OpenSmartRoute, default objective (cost weight 0.15) | 0.716 | 0.682 | $0.34 |
| Random | 0.552 | 0.521 | $0.84 |
| Always the cheapest model (mistral-7b) | 0.317 | 0.310 | $0.05 |
What the router buys is the cost axis. At the default objective it spends $0.34 per thousand calls against $2.98 for the table and $3.31 for always-frontier, for eleven points of accuracy. Told that quality is all that matters, it reaches 0.810 with a higher realised quality than the table.
Where it loses. With real task labels a lookup table fitted on the training half is a strong, cheap competitor: 0.830. Between cost weights 0.1 and 0.3 the frontier drops off a cliff, because on this catalogue the cheap models are much worse - the operator chooses the side of the cliff; the router does not pretend there is a middle.
Calibration. Raw confidence with eleven candidates was badly under-confident (ECE 0.55); a calibrator fitted on the even rows brings the odd rows to ECE 0.01. The platform ships one with every catalogue this wide.
Public preference suites
Routing accuracy on 1,000 prompts per suite: which of three model tiers was good enough. The learned router against the same catalogue without learning, a static task table, random choice and always-cheapest.
The learned router beats the declarative one by 9 to 38 points and random and cheapest everywhere, and does not beat the best single tier - a pairwise battle between two named models is a weak label for a three-tier decision. The results page says so in the same words.
| Suite | Router | No learning | Static table | Random | Cheapest |
|---|---|---|---|---|---|
| MT-Bench humanlmsys/mt_bench_human_judgments | 0.514 | 0.421 | 0.541 | 0.441 | 0.372 |
| PPE humanlmarena-ai/PPE-Human-Preference-V1 | 0.563 | 0.418 | 0.651 | 0.445 | 0.174 |
| WebDev Arenalmarena-ai/webdev-arena-preference-10k | 0.573 | 0.192 | 0.670 | 0.391 | 0.000 |
Reproduce it
The multi-cloud run calls real providers on your own accounts and prints the same tables; the RouterBench run needs only the public dataset. Results are deterministic for the same seed and cache.
pip install "opensmartroute[yaml]"
# the 150-prompt multi-cloud run (real calls on your own provider accounts)
python examples/multicloud_benchmark.py --prompts examples/multicloud_prompts.jsonl \
--rules examples/multicloud_rules.yaml --judge azure-gpt-4.1,gcp-gemini-2.5-pro
# RouterBench 0-shot, honest split, about 15 minutes after the download
hf download withmartian/routerbench --repo-type dataset --local-dir rb
python examples/leaderboard/routerbench.py --pkl rb/routerbench_0shot.pkl --work scratchYour own traffic is the benchmark that matters: upload last month's provider invoice and the bill analysis prices every line through the router, or run osr eval on a labelled sample of your prompts.
The free plan keeps the full decision for every request - the candidates, the policy rejections, the cost of each - so you can judge the router on your traffic, not ours.
Free plan, no card. Fifteen thousand decisions a month with the full trace.