Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
CoreWeave launches Forge platform for AI inference optimization - OpenSmartRoute
CoreWeave launches the Forge platform at the Fully Connected event
CoreWeave Inc. unveiled a new platform called Forge. This launch happened at the Fully Connected event in late October 2026. Urvashi Chowdhary, vice president of product and AI services, announced it during an exclusive broadcast. The event took place on SiliconANGLE Media's livestreaming studio. Dave Vellante and John Furrier from theCUBE Research joined her for the interview. They discussed how CoreWeave is changing the AI infrastructure landscape.
CoreWeave targets specific bottlenecks in AI inference workloads. Inference demand is rising faster than training demand in many cases. A survey of customers showed one healthcare client's inference share grew to 40% in its second year. Experts expect that figure to reach about 50% within the next 12 months. CoreWeave focuses on tuning every layer above the raw hardware. This includes the vLLM engine, quantized models, and custom speculative decoders.
The company is moving beyond just providing GPU capacity. They are building managed services for training, post-training, and inference. Chowdhary explained that developers want to solve problems quickly with good performance. They also want to control the cost to scale their applications. CoreWeave aims to build layers of its stack on top of reliable infrastructure. This approach allows them to offer more managed services than before.
CoreWeave Forge connects serving, observability, post-training, and evaluation tools. It is designed to optimize the entire AI inference stack layer by layer. The platform is free to start using for individual developers. Paid tiers exist for teams that need additional capabilities or extended access. These tiers offer more development tools and services beyond the basic infrastructure.
Chowdhary emphasized creating accessibility for leading technologies in the field. Individual developers can sign up today without compromising on performance. They get the best reliability available from CoreWeave's infrastructure. The platform ensures they do not have to sacrifice quality for speed or cost. This openness is a key part of their strategy for the market.
CoreWeave Forge represents a shift in how cloud providers serve AI models. Traditional clouds focused heavily on training workloads with massive GPU clusters. Now, serving models faster and cheaper defines the next wave of competition. Inference has become the workload that decides the economics of the AI boom. Providers are pushing into storage, networking, and software to support this shift.
OpenAI released 719 math manuscripts from an unreleased frontier model. Mistral and Reflection AI launched new open-weight models to compete with Chinese systems.
Perplexity launched two ColBERT-style embedding models on Hugging Face under an MIT license.
The Fully Connected event featured an exclusive broadcast on theCUBE. This was part of SiliconANGLE Media's coverage of the latest trends in AI. The discussion highlighted CoreWeave RL Rollouts as a new capability for agentic models. This feature is built on Nvidia Corp.'s Dynamo framework technology. It allows teams to load new checkpoints into a live deployment quickly.
CoreWeave RL Rollouts specifically addresses reinforcement learning challenges. When customers train agentic models with rewards and verifiers, inference becomes the bottleneck. CoreWeave's solution speeds up the process of creating new model checkpoints. These checkpoints then roll out into the inference setup for independent scaling.
Chowdhary noted that training runs fast when inference scales while continuing to train. The capability improved model reload latency by 15x compared with a baseline configuration. This means teams can iterate on their models much faster during deployment. It supports the continuous creation of new versions without significant delays.
CoreWeave Forge integrates these capabilities into a unified platform. It serves as a hub for managing the full lifecycle of AI models. Teams can now handle serving, monitoring, and evaluation from one place. This reduces the complexity of managing multiple disparate tools. The platform aims to streamline the journey from training to production deployment.
The shift from training-focused clouds to inference-focused services is reshaping the industry. Training built the first wave of GPU clouds with massive compute power. Serving models faster and cheaper will define the next generation of cloud providers. Inference demand is climbing fast across various sectors including healthcare. CoreWeave's answer involves tuning every layer above the hardware.
Optimizing the AI inference stack requires attention to detail at every level. From the vLLM engine to quantized models, each component affects performance. Custom speculative decoders help reduce the time needed for model responses. Open source tools and technologies play a crucial role in this optimization. CoreWeave is intentional about leveraging these open systems.
Contributing back to open systems gives customers flexibility in their choices. Building services on top of each other creates a cohesive ecosystem. This approach ensures that customers have access to the latest innovations. It also fosters collaboration between different parts of the AI development process.
Reinforcement learning adds new pressure to the existing infrastructure demands. Agentic models require continuous iteration and fine-tuning based on rewards. Inference becomes the bottleneck during these rollouts, slowing down the entire process. CoreWeave RL Rollouts is designed to alleviate this specific challenge.
CoreWeave Forge connects serving, observability, post-training, and evaluation tools. This integration allows for a more streamlined development workflow. Teams can monitor their models in real-time while they are being trained. Post-training services help refine the model before it goes live. Evaluation ensures the model meets quality standards before deployment.
The platform is free to start, making it accessible to smaller teams. Paid tiers offer additional capabilities for larger organizations or specific needs. This structure allows individual developers to benefit from the technology immediately. It also provides a path for scaling up as their needs grow.
Chowdhary stated that even individual developers get the best performance available. They do not have to compromise on reliability or quality. The platform supports teams of all sizes with its tiered approach. This inclusivity is a key factor in CoreWeave's market strategy.
CoreWeave Forge represents a significant step forward for AI infrastructure. It addresses the growing complexity of managing modern AI workloads. By connecting multiple layers of the stack, it reduces operational overhead. Teams can focus more on model development and less on infrastructure management.
The shift towards inference-focused services is driven by market demand. Training costs are high, but serving costs are becoming equally critical. Inference demand is climbing fast in sectors like healthcare and finance. CoreWeave's platform aims to meet this growing demand efficiently.
CoreWeave RL Rollouts accelerates reinforcement learning rollouts significantly. It loads new checkpoints into a live deployment with minimal disruption. The 15x improvement in model reload latency is a concrete metric of its success. This speed allows for faster iteration cycles during the training phase.
When doing RL rollouts, teams create new model checkpoints and versions constantly. They want these to roll out into their inference setup independently. CoreWeave's solution enables this independent scaling without waiting for full retraining. It keeps the training process moving while the model is in use.
CoreWeave Forge connects serving, observability, post-training, and evaluation tools. This integration creates a unified view of the AI development lifecycle. Teams can track performance from the first checkpoint to production deployment. Observability helps identify issues before they impact users.
The platform is free to start, which lowers the barrier to entry. Paid tiers exist for teams that need advanced features or higher limits. This model allows small teams to try it out without financial risk. It also provides a clear path for scaling up as needed.
Chowdhary explained that the goal is to create accessibility for leading technologies. Individual developers can sign up today and get the best performance. They do not have to compromise on reliability or quality. The platform ensures they have access to top-tier infrastructure.
CoreWeave Forge represents a shift from providing raw GPU capacity to full-stack optimization. It addresses the specific bottlenecks that slow down AI inference workloads. By tuning every layer, it maximizes the efficiency of the hardware. This approach is essential for managing the economics of the AI boom.
The Fully Connected event highlighted CoreWeave's commitment to innovation in AI infrastructure. The exclusive broadcast on theCUBE brought together industry leaders to discuss these changes. Urvashi Chowdhary shared insights into their strategy and roadmap. Her comments focused on solving problems quickly with good performance and cost.
CoreWeave targets AI inference bottlenecks with full-stack optimization. This focus distinguishes them from providers that only offer raw compute. They are layering managed services for training, post-training, and inference. This comprehensive approach addresses the needs of developers at every stage.
The survey by theCUBE Research found a significant rise in inference workload share. One healthcare customer's share rose from about 10% to 40% in two years. A roughly 50% share is expected within the next 12 months. This trend underscores the growing importance of efficient inference services.
CoreWeave's answer involves tuning every layer above the hardware. From the vLLM engine to quantized models and custom speculative decoders, optimization is key. They are leveraging open source tools and technologies to achieve this. Contributing back to open systems ensures customer flexibility and innovation.
Reinforcement learning adds new pressure to the existing infrastructure demands. Inference becomes the bottleneck during rollouts when training agentic models with rewards. CoreWeave RL Rollouts is a preview capability built on Nvidia Corp.'s Dynamo framework. It improves model reload latency by 15x compared with a baseline configuration.
CoreWeave Forge integrates these capabilities into a unified platform for serving, observability, post-training, and evaluation. It is free to start, with paid tiers offering additional capabilities. This structure supports individual developers and larger teams alike. The goal is to create accessibility for leading technologies without compromise.
Why full-stack optimization matters for AI developers and operators
Full-stack optimization reduces the time it takes to build and deploy AI models. Teams spend less time managing infrastructure and more time developing their applications. Inference costs can drop significantly when every layer is tuned correctly. This efficiency directly impacts the economics of running an AI business.
Developers face complex challenges when moving from training to serving models. Training builds the first wave of GPU clouds, but serving defines the next. Inference demand is climbing fast across various industries including healthcare. Without optimization, these demands can lead to higher costs and slower performance.
CoreWeave Forge addresses these challenges by connecting multiple layers of the stack. It serves as a hub for managing the full lifecycle of AI models. Teams can monitor their models in real-time while they are being trained. Post-training services help refine the model before it goes live. Evaluation ensures the model meets quality standards before deployment.
The shift towards inference-focused services is driven by market demand for speed and cost. Training costs are high, but serving costs are becoming equally critical. Inference demand is climbing fast in sectors like healthcare and finance. CoreWeave's platform aims to meet this growing demand efficiently through optimization.
CoreWeave RL Rollouts accelerates reinforcement learning rollouts significantly. It loads new checkpoints into a live deployment with minimal disruption. The 15x improvement in model reload latency is a concrete metric of its success. This speed allows for faster iteration cycles during the training phase.
When doing RL rollouts, teams create new model checkpoints and versions constantly. They want these to roll out into their inference setup independently. CoreWeave's solution enables this independent scaling without waiting for full retraining. It keeps the training process moving while the model is in use.
CoreWeave Forge integrates serving, observability, post-training, and evaluation tools. This integration creates a unified view of the AI development lifecycle. Teams can track performance from the first checkpoint to production deployment. Observability helps identify issues before they impact users.
The platform is free to start, which lowers the barrier to entry for many teams. Paid tiers exist for organizations that need advanced features or higher limits. This model allows small teams to try it out without financial risk. It also provides a clear path for scaling up as needed.
Chowdhary explained that the goal is to create accessibility for leading technologies in the field. Individual developers can sign up today and get the best performance available. They do not have to compromise on reliability or quality. The platform ensures they have access to top-tier infrastructure.
CoreWeave Forge represents a shift from providing raw GPU capacity to full-stack optimization. It addresses the specific bottlenecks that slow down AI inference workloads. By tuning every layer, it maximizes the efficiency of the hardware. This approach is essential for managing the economics of the AI boom.
The Fully Connected event highlighted CoreWeave's commitment to innovation in AI infrastructure. The exclusive broadcast on theCUBE brought together industry leaders to discuss these changes. Urvashi Chowdhary shared insights into their strategy and roadmap. Her comments focused on solving problems quickly with good performance and cost.
CoreWeave targets AI inference bottlenecks with full-stack optimization. This focus distinguishes them from providers that only offer raw compute. They are layering managed services for training, post-training, and inference. This comprehensive approach addresses the needs of developers at every stage.
The survey by theCUBE Research found a significant rise in inference workload share. One healthcare customer's share rose from about 10% to 40% in two years. A roughly 50% share is expected within the next 12 months. This trend underscores the growing importance of efficient inference services.
CoreWeave's answer involves tuning every layer above the hardware. From the vLLM engine to quantized models and custom speculative decoders, optimization is key. They are leveraging open source tools and technologies to achieve this. Contributing back to open systems ensures customer flexibility and innovation.
Reinforcement learning adds new pressure to the existing infrastructure demands. Inference becomes the bottleneck during rollouts when training agentic models with rewards. CoreWeave RL Rollouts is a preview capability built on Nvidia Corp.'s Dynamo framework. It improves model reload latency by 15x compared with a baseline configuration.
CoreWeave Forge integrates these capabilities into a unified platform for serving, observability, post-training, and evaluation. It is free to start, with paid tiers offering additional capabilities. This structure supports individual developers and larger teams alike. The goal is to create accessibility for leading technologies without compromise.
How OpenSmartRoute helps teams manage CoreWeave's new capabilities
Teams that route their requests through OpenSmartRoute gain better control over AI spending. They can set rules so personal data stays on an on-premises model. A region or a cost cap is never crossed by any routed request. The input guard spots prompt injection and personal data before it leaves.
OpenSmartRoute scores every candidate on quality, cost, speed, and safety. The team sets the weights per request to prioritize what matters most. It learns from outcomes, so a model that answers well gets more traffic. One that fails gets less traffic automatically over time.
A new model is one catalogue entry in OpenSmartRoute's system. It competes on the next request; nothing else changes in the app. The hosted platform keeps a models catalogue with prices and public rankings. These rankings are built from real traffic data collected by the router.
The savings ledger shows what each routed request cost next to what the most expensive model would have cost. This transparency helps teams understand their total infrastructure expenses clearly. It highlights where optimization efforts are saving money compared to baseline spending.
CoreWeave Forge connects serving, observability, post-training, and evaluation tools. OpenSmartRoute can integrate with these tools to manage the full lifecycle of models. Teams can use CoreWeave's RL Rollouts while OpenSmartRoute routes traffic efficiently. This combination maximizes both performance and cost efficiency.
OpenSmartRoute works with any OpenAI-compatible provider, including CoreWeave's infrastructure. It supports open-weight models served locally as well as cloud-based options. MCP tools and A2A agents are also supported for complex workflows. This flexibility allows teams to mix and match the best components.
The 'osr eval' command measures routing accuracy on the team's own prompts. It can fail a build when it drops below a certain threshold. This ensures that the routing logic remains accurate and reliable over time. Teams can test their setup before deploying it to production traffic.
CoreWeave Forge represents a shift towards inference-focused services in the cloud market. OpenSmartRoute helps teams navigate this shift by optimizing how they spend on inference. It provides the control needed to manage costs while maintaining high performance. Together, they offer a powerful solution for modern AI development.
What to do
Check CoreWeave's documentation to understand how Forge connects serving and evaluation tools. Look at their pricing tiers to see if the free start is enough for your team. Test CoreWeave RL Rollouts with a small project to measure the 15x latency improvement.
Set up OpenSmartRoute to route traffic through CoreWeave while enforcing cost caps. Define hard rules so personal data stays on-premises if required by your policy. Use 'osr eval' to verify that your routing accuracy meets your team's standards.
Compare the savings ledger from OpenSmartRoute against your current spending on inference services. Identify which models are costing more than necessary and adjust their weights accordingly. Consider using CoreWeave's post-training services to refine models before they enter your catalogue.
Explore how open-weight models served locally can reduce costs for sensitive data tasks. Ensure your input guard is active to prevent prompt injection attacks from reaching the model. Combine these tools with CoreWeave's full-stack optimization to build a resilient AI system.
How OpenSmartRoute helps
For a team that routes through the router, a release like this is one catalogue entry. The new model competes on the next request against what the team already runs, on quality, cost and speed. If it answers well it earns more traffic; if it fails it gets less, with no code change in the app.
Any OpenAI-compatible provider works, as do open-weight models served locally. The catalogue at https://opensmartroute.ai lists models with their prices and live rankings, and the savings ledger shows what the routing saved against always using the most expensive option.