Glossary
An LLM router is a component that sits between an application and several large language models and decides, for each request, which model should answer it - by reading the request and scoring the candidates on quality for that task, cost, latency and the caller's rules.
Applications that use more than one language model have to choose between them somewhere. Without a router that choice is a configuration value or a condition in the code - a model name per feature, an if/else on prompt length - and it stays fixed while the models, their prices and their quality change around it. An LLM router moves the choice out of the application into a layer that makes it per request, from the request itself.
A router works in three steps. It reads the request and extracts signals: the task (summarise, translate, write code, answer a question), its complexity, its domain and language, whether it carries personal data or code. It then scores each candidate model on how well it handles that task - from published benchmarks, declared capabilities, a small routing model and the outcomes the caller has reported - together with what the request would cost at this length and how long the model takes. The caller's rules bound the choice: a hard constraint such as a data boundary removes a candidate outright, a preference adjusts its score. The highest-scoring candidate answers.
A good router explains itself. Every decision records the alternatives, their scores and costs, and the rules that removed or favoured a candidate, so a developer can see why a request went where it went and a finance lead can read what the same traffic would have cost on the most expensive model alone. It also learns: when the caller reports whether an answer was good, the next decision for a similar request moves.
An LLM router is not a proxy or a gateway, though it is usually deployed as one. A proxy forwards a request to the model the caller named; a gateway adds keys, rate limits, retries and logging; a router adds the decision. The three are complementary, and most routers - this one included - expose an OpenAI-compatible endpoint so an existing client can be pointed at them without a code change, with `model: "auto"` meaning "you decide".
Questions people ask
OpenSmartRoute is open source and the free plan keeps the full trace of every decision. Type a request in the playground and read the ranked candidates.
Free plan, no card. Fifteen thousand decisions a month with the full trace.