Glossary
Model fallback is what a router does when the chosen model fails, times out or is saturated: it re-routes the request to the next best candidate under the same rules instead of returning the error, and records both the failure and the recovery.
Providers fail: a rate limit, a timeout, an outage, a model deprecated overnight. An application that retries the same model waits out the problem; one that hard-codes a second model is right until that one changes too. Fallback in a router means the request goes to the next best candidate from the same ranking that chose the first - a different model, perhaps a different vendor - and the application sees an answer rather than an error.
Fallback has to respect policy. A request pinned to an on-premises model by a data-boundary rule must not fall back to a public one because the on-premises model is down; it should queue, degrade or fail loudly. Because the router applies constraints before scoring, the fallback candidates are already the allowed ones. Per-target circuit breakers stop sending traffic to a provider that keeps failing and let it back in gradually, so one bad provider does not slow every request.
The trace should show it. A recovered request records the target that failed, the error, the target that answered and the cost of both, and the comparison is fed back to the learners as a preference - the target that answered beat the one that did not. Over time a flaky provider earns a lower score on its own evidence.
Questions people ask
OpenSmartRoute is open source and the free plan keeps the full trace of every decision. Type a request in the playground and read the ranked candidates.
Free plan, no card. Fifteen thousand decisions a month with the full trace.