Self-hosted LLM Gateway: Open Source Model Serving Guide
Engineers need a reliable self-hosted LLM gateway to manage private model inference without relying on public APIs. OpenSmartRoute provides the infrastructure to deploy inference gateways across various environments, from single VMs to Kubernetes clusters, ensuring control over data and costs.
Deployment Shapes and Environments
The choice of deployment shape depends on the scale and the underlying infrastructure. For a laptop or a single virtual machine, the osr-models binary or a Docker run command offers an Ollama-equivalent experience. This setup uses loopback networking and commands like osr-models pull, osr-models run, and osr-models ps to manage the service. In Docker Compose environments, the architecture supports one osr-models service per profile, such as llm, gpt-oss, or voice, with the gateway acting as a service and the controller residing inside the gateway container. Azure Container Apps also support this pattern, where the gateway becomes an internal application that the API calls, while the controller scales revisions through the ARM API.
For enterprise-scale deployments, Kubernetes offers a Helm chart approach. This includes a Deployment per pool for CPU resources or a StatefulSet for GPU requests. The architecture utilizes an InferencePool and the Enterprise Platform Provider (EPP), along with a KEDA ScaledObject based on gateway metrics. Persistent volumes or OCI volumes can store the model store. Bare metal or on-premises installations utilize systemd units from the installer, with an optional shared NFS store. In the enterprise edition's air-gapped case, a registry mirror is placed within the same site.
Gateway Architecture and Control
The distinction between the router and the gateway is critical for understanding the system. The router produces a target, such as llm-onprem, and the provider handler resolves this to a model instance group, for example, qwen3.5:4b on a cpu-small pool. The gateway then produces a replica, specifying a node, instance, and slot. Strategies that require gateway signals, such as expected queue wait times or prefix hit probabilities per pool, read these metrics through the provider's metadata. The calibrator refreshes this metadata, and nothing within needs to know the gateway exists directly.



