LLM cost optimization
Route each request to the cheapest model that meets your quality rules, keep the frontier model for the requests that need it, and read the saving per request, key, tenant and month. Upload last month's invoice first and see the number before you switch.
Free plan, no card. Open source under Apache-2.0; self-hosting is free forever.
From the benchmark
$176 vs $1,638
per million requests, routed vs always-frontier
On 150 real prompts, projected to a million requests of the same mix: the router at $176 against $1,638 for the frontier model alone - with the routed answers judged higher on average.
The whole benchmark, with the cases the router lostThe problem
Most requests to a frontier model would have been answered as well by one that costs a tenth. Nobody can tell which ones from the invoice, so the whole application pays the top rate - and finance sees a line that doubles every quarter with no lever attached.
How the router answers it
Upload the invoice PDF or usage export from your provider. Every line is priced twice - what you paid and what the router would have charged for the same traffic. That is the pilot.
Your rules say which model may answer what; within them the cheapest wins. Budgets per workspace, tenant and key are ceilings that refuse, not alerts that arrive later.
The savings ledger records the baseline and the routed cost for every executed request, by key, tenant and month - exportable for the person who signs.
A real request
The analysis runs on the free plan, stays in your workspace and can be deleted. The figures are estimates and the report says so; the live ledger replaces them once traffic flows.
date,model,input_tokens,output_tokens,cost
2026-08-01,claude-sonnet-4,120000,30000,0.81
2026-08-01,gpt-4.1,95000,22000,0.37
2026-08-02,gpt-4o,210000,48000,1.01
# -> every line priced through the router; CSV, workbook or PDF reportRead next
Install
osr audit traffic.jsonl --baseline gpt-4.1 --markdownOr no install at all: the hosted API answers a plain curl with the decision, and the browser extension shows the price and the router's pick under the composer of the AI chat sites you already use.
OpenSmartRoute is open source and the free plan keeps the full trace of every decision. Create a workspace, mint a key and send the request above.
Free plan, no card. Fifteen thousand decisions a month with the full trace.