Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Prime Intellect launches Prime Inference for open models with serverless and reserved options - OpenSmartRoute
Introduction - Prime Inference launches as a new serving platform for open models
Prime Intellect has introduced Prime Inference, a new platform for serving open-source machine learning models. It is designed to handle large models used in AI research and deployment. The platform aims to make it easier to run these models at scale. It supports both flexible, on-demand use and dedicated, reserved resources.
Prime Inference is part of Prime Intellect’s broader open training stack. This stack includes tools for training, evaluating, and deploying models. Prime Inference specifically focuses on the serving layer, which runs models in production. It is built to support models used in agentic tasks, such as chatbots or coding assistants. The platform is designed to work with models from the frontier open-source community.
Before its public release, Prime Inference processed nearly a trillion tokens per day internally. Tokens are pieces of text, and processing this many shows the platform’s capacity. The traffic came from various tasks, including reinforcement learning rollouts, synthetic data generation, and evaluation. It also supported long-running coding agents, which are AI systems that write or debug code over extended sessions. This high internal volume indicates the platform’s readiness for real-world workloads.
Prime Inference is compatible with open models and aims to serve them efficiently. It offers features that help manage demand, ensure uptime, and optimize performance. The platform is designed to support the needs of AI developers and researchers who want to deploy models quickly and reliably. Its launch marks a step toward more accessible and scalable open-model serving.
Key features - Serverless endpoints, reserved capacity, and multi-datacenter failover
Prime Inference offers serverless endpoints, which allow users to run models without managing servers. These endpoints automatically scale with demand. When many users request predictions, the platform adds resources. When demand drops, it reduces resources. This flexibility helps control costs and handle variable workloads.
For workloads that are steady or predictable, Prime Inference also provides reserved capacity options. Users can reserve dedicated GPU resources in advance. This ensures consistent performance and availability for long-term projects. Reserved capacity is suitable for applications with steady usage or critical production needs.
Reflection released Beam, an open-weight model that matches GLM 5.2 and Qwen 3.8 on benchmarks while using three to four times less compute.
The platform supports automatic failover across multiple datacenters. If one datacenter experiences issues, traffic is rerouted to healthy deployments in other locations. This setup improves reliability and uptime. It helps prevent service interruptions, even during hardware failures or network problems. The failover system is designed to keep models available at all times.
Prime Inference’s combination of serverless and reserved options gives users flexibility. They can choose the best deployment mode for their workload. The multi-datacenter failover adds an extra layer of reliability. These features help ensure that open models run smoothly and continuously.
Prime Inference runs on high-performance hardware, starting with NVIDIA Blackwell GPUs. These GPUs are designed for AI workloads, offering high speed and efficiency. The platform also plans to support Vera Rubin GPUs soon, which are expected to provide additional capacity.
The software stack includes several open-source components. NVIDIA Dynamo handles routing of requests and manages GPU resources. vLLM is used to run models on GPU groups efficiently. Mooncake provides a second tier of cache in host memory, which helps speed up data access. FlashInfer is used for fast inference and optimized model execution.
This stack was built with contributions from NVIDIA and the open-source community. It combines these tools to create a flexible, scalable serving platform. The software is designed to fix issues and improve performance through upstream contributions. This approach helps keep the platform up to date with the latest advances in AI hardware and software.
Prime Inference supports models that run on NVIDIA Blackwell GPUs today. The hardware and software choices aim to maximize throughput and minimize latency. They also support large models with high memory needs. The platform’s architecture is designed to adapt as new hardware and software tools become available.
Performance benchmarks - Latency, throughput, and capacity metrics
Prime Inference has been tested to measure how fast and how many predictions it can handle. In benchmarks, the platform achieved low latency, which means predictions are delivered quickly. Specifically, tests show a nearly 40% reduction in the worst-case (p90) inter-token latency compared to previous setups. This means responses come faster, especially in complex tasks.
Throughput, or the number of predictions per second, is also high. The platform can serve up to 66 sessions per prefill group at a rate of about 101 tokens per second per user. Each session involves multiple turns of interaction, with the system processing thousands of tokens. The setup can handle 100 output tokens per GPU per second, supporting interactive applications.
Capacity is improved through architectural innovations. Prime Inference increased the number of cached tokens per decoder from about 1.09 million to 1.63 million. This allows models to process larger contexts and reduces the need to fetch data repeatedly. The platform also uses compression techniques to shrink transfer descriptors, reducing data movement times significantly.
These performance metrics demonstrate the platform’s ability to support demanding workloads. It can serve many users simultaneously with low delays. The improvements help make open models more practical for real-time applications.
Workload focus - Agentic tasks with specific token and session targets
Prime Inference is optimized for agentic workloads. These are tasks where AI models act as agents, such as chatbots or coding assistants. Such workloads typically involve many interactions, each adding tokens to the conversation or task context.
A typical agent turn adds about 6,000 tokens to a prompt of 140,000 tokens. The platform benchmarks this workload using tools like SemiAnalysis AgentX. It also tests with injected cold arrivals, which simulate sudden increases in demand. These tests help measure how well the system handles real-world usage patterns.
The platform aims to support interactive applications with high responsiveness. It targets processing 100 tokens per second per user, which is suitable for conversational AI. The setup can serve multiple sessions simultaneously, with a prefill/decode ratio of about 1:4. This ratio balances the initial data load with ongoing predictions.
Prime Inference’s focus on agentic workloads makes it suitable for AI systems that require continuous, high-speed interaction. Its architecture supports large contexts and quick responses, essential for user-facing AI applications.
Technical innovations - Prefill/decode disaggregation, cache-aware routing, and KV compression
Prime Inference introduces several technical innovations to improve performance. One key feature is prefill and decode disaggregation. Prefill runs on one GPU group, preparing the initial context. Decode runs on another group, generating tokens. Dynamo manages routing between these groups, reducing delays.
This separation allows each part to be optimized independently. Prefill can be scaled separately from decoding, improving overall throughput. It also reduces queue wait times. For example, halving the tokens per step from 8,000 to 4,000 decreased median queue wait from 550 milliseconds to 110 milliseconds.
Cache-aware routing is another innovation. Dynamo’s router considers cached prefix overlap when assigning work. If a session’s context is already cached, the system avoids redundant data fetches. Mooncake adds a second cache tier in host memory, further speeding up data access.
Prime Inference also uses compression techniques. NVFP4 KV compression shrinks cache rows from 576 bytes to 352 bytes. This allows more tokens to be stored in cache, increasing the number of cached tokens per decoder from about 1.09 million to 1.63 million. These innovations reduce latency and increase capacity.
The platform also employs a native sparse-MLA kernel, which performs specific workloads faster. For example, it achieves about 12 microseconds per 15 query tokens, compared to staged or FP8 kernels. These technical improvements help support large, fast models.
Safety and reliability - Failover, error rates, and tool call handling improvements
Prime Inference emphasizes safety and reliability in its design. The platform supports automatic failover across multiple datacenters. If one location encounters issues, traffic is rerouted to healthy deployments. This setup helps maintain high uptime and service continuity.
Error rates are kept very low. The platform reports a near-zero tool-call error rate, meaning that AI models can call tools or APIs without frequent failures. This stability is crucial for agentic tasks that depend on external tools.
Prime Inference also improves handling of tool calls. It contributed a structural-tag builder to Dynamo, which helps models format tool requests correctly. vLLM uses xgrammar to mask tokens that violate tool schemas, preventing errors. These measures help ensure that models interact safely and correctly with external tools.
The focus on safety and reliability helps users deploy models in production environments. It reduces the risk of unexpected failures and improves the overall quality of AI interactions.
Pricing and deployment options - On-demand and reserved capacity, billing, and upcoming features
Prime Inference offers flexible deployment options. Users can choose on-demand dedicated GPU deployments for variable workloads. These are billed based on usage, allowing for cost control. The platform also plans to add one-click dedicated deploys, simplifying the process of reserving resources.
Reserved capacity options are available for steady workloads. Users can reserve GPUs in advance, ensuring consistent performance. Prime Intellect plans to introduce one-click dedicated deploys for reserved capacity as well. This will make long-term deployment easier and more predictable.
Billing is unified across the platform, with team-level usage tracking. This helps organizations monitor costs and optimize their deployments. Per-model pricing details are not yet fully published, but the platform provides transparency in overall usage.
Upcoming features include batch inference, which allows processing multiple inputs at once, and more deployment options. These features aim to make deploying open models more flexible and cost-effective.
Why it matters and what to do - Benefits for model deployment and next steps for users
Prime Inference makes deploying open models easier and more reliable. Its flexible options help match workload demands with the right resources. The platform’s technical innovations improve speed, capacity, and stability.
For model developers and users, Prime Inference offers a way to run large models at scale without managing hardware. Its serverless endpoints reduce costs for variable workloads. Reserved capacity ensures performance for steady, long-term use.
The platform’s focus on safety and reliability helps users trust their models in production. Its multi-datacenter failover protects against outages. These features support continuous, high-quality AI services.
Users can check the technical details and documentation to understand how to deploy models on Prime Inference. They can also explore upcoming features like batch inference and dedicated deploys. The platform aims to simplify open-model deployment and improve operational efficiency.
Organizations interested in using Prime Inference should evaluate their workload patterns. They can start with on-demand deployment and consider reserving capacity for steady use. Monitoring usage and performance will help optimize costs and reliability.