Why a router
Every application that calls more than one model grows an if/else that picks between them - and it is wrong a month later, because models, prices, latency, quality, providers and policies all change. An AI router moves that decision into one layer that reads each request, applies your policy and keeps up. Here is the argument, the architecture and the way in.
Free plan, no card. Open source under Apache-2.0; self-hosting is free forever.
From the benchmark
89%
less spent than always using the frontier model
On 150 real prompts the router cost 89% less than sending every one to the frontier model, with answers two blind judges scored higher (9.53 vs 9.15). The methodology, the prompts and the cases the router lost are on the benchmark page.
The whole benchmark, with the cases the router lostThe problem
A setting, a condition on prompt length, a person who 'knows which one is good at code'. Then a cheaper model ships, a provider has an outage, a price doubles, a team pastes customer data into the expensive one, and finance asks why the bill went up. Each answer is a code change, in every application, with no record of why a request went where it went.
How the router answers it
The application sends every request to one OpenAI-compatible URL. Which model - or agent, tool or person - answers is decided there, per request, from the request itself.
Data boundary, region, tenant and budget remove what may not answer; what remains is ranked on quality for this task, cost for this length and latency. The lowest-cost option that satisfies the rules wins.
Every decision carries the alternatives, the cost and the reason. Report whether the answer was good and the next decision moves - without a deploy.
A real request
A support reply routes to a small model with the support persona; the same call with a contract analysis picks a reasoning model, and with 'I want to speak to a person' a human queue. The application code is identical.
from openai import OpenAI
client = OpenAI(base_url="https://opensmartroute.ai/v1",
api_key=OSR_API_KEY)
reply = client.chat.completions.create(
model="auto", # the router picks
messages=[{"role": "user", "content": "<prompt>"}],
)
print(reply.choices[0].message.content)
print(reply.model) # the model that answeredRead next
Install
pip install "opensmartroute[yaml]"Or no install at all: the hosted API answers a plain curl with the decision, and the browser extension shows the price and the router's pick under the composer of the AI chat sites you already use.
OpenSmartRoute is open source and the free plan keeps the full trace of every decision. Create a workspace, mint a key and send the request above.
Free plan, no card. Fifteen thousand decisions a month with the full trace.