Glossary
LLM cost is what an application pays to call language models, almost always priced per million input and output tokens with rates that differ by one to two orders of magnitude between models; LLM cost optimisation lowers it by sending each request to the cheapest model that meets the quality required.
Language models are priced per token - a few characters of text - with separate rates for input and output and, often, a discount for cached input. The rates span a wide range: the most capable models cost ten to a hundred times more per token than small ones. An application's bill is therefore its traffic volume times the prices of the models it chose, and most applications chose once.
Most of the bill buys quality that the request did not need. A frontier model answers a one-line rewrite and a forty-page analysis at the same rate; a small model answers the first as well and the second badly. The largest single lever is to route each request to the cheapest model that meets the quality it needs, and keep the frontier model for the requests that need it. Prompt size, caching and output limits help at the margin; routing changes the rate.
The saving has to be measured, not asserted. Two numbers do it: the routed cost of real traffic and what the same traffic would have cost on the most expensive model alone, both from list prices, with the judged quality alongside so a cheaper answer that was worse is counted against the saving. A savings ledger records both for every request; a bill analysis prices last month's invoice both ways before anything is switched.
Questions people ask
OpenSmartRoute is open source and the free plan keeps the full trace of every decision. Type a request in the playground and read the ranked candidates.
Free plan, no card. Fifteen thousand decisions a month with the full trace.