Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Reflection's Beam beats Chinese models with less compute - OpenSmartRoute
Reflection's Beam beats Chinese models with less compute
Reflection released Beam, an open-weight model that matches GLM 5.2 and Qwen 3.8 on benchmarks while using three to four times less compute.
Key points
Beam activates only 23 billion of its 501 billion parameters per token.
It uses three to four times less compute than GLM 5.2 on reasoning tasks.
The model was trained on 10,500 Nvidia GPUs for over four weeks.
Beam matches Qwen 3.8-Max on coding and agent benchmarks.
Why it matters: Businesses can run strong AI agents for coding with lower operating costs.
By OpenSmartRoute editorial · written through the router by writer-small
From The Decoder - “Reflection's Beam becomes the most capable open-weight model built outside China”
An abstract graphic consisting of white dots and lines on a dark green background, spreading from a small circle on the left into increasingly larger rings toward the right.. Image: The Decoder (original)
Reflection releases Beam as its first open-weight model
The AI startup Reflection is launching a new model called Beam. This release marks the company's entry into the open-weight market. Beam focuses on coding and reasoning tasks while prioritizing efficiency. The team built this model to compete with Chinese alternatives like Deepseek and Qwen.
Beam aims to deliver strong results without requiring massive compute resources. Engineers can run this model locally or in the cloud with lower costs. The company states that Beam activates only a fraction of its total parameters. This design choice keeps the energy required for inference significantly down.
The model targets businesses needing AI for automated workflows and software development. Traditional open models often demand huge data centers to run effectively. Beam challenges this norm by proving high performance is possible with less power. It positions itself as a Western alternative to major Chinese research groups.
Architecture - How the mixture-of-experts design saves compute
Beam uses a specialized architecture known as a mixture of experts. This system activates only 23 billion parameters for every single token generated. The total model size reaches 501 billion parameters, but most remain dormant.
Activating fewer parameters reduces the computational load during inference. Standard models usually turn on all their weights to generate text. Beam's approach allows it to match performance while using three to four times less compute. This efficiency comes from selectively engaging only the necessary parts of the network.
The architecture supports a tunable parameter that controls reasoning depth. Users can choose between quick answers or deeper thought processes. Selecting more time for thinking increases output quality on difficult tasks. Conversely, choosing speed lowers the cost per generated token.
This flexibility makes Beam suitable for various production environments. Teams can adjust settings based on their specific budget and deadline constraints. The design ensures that compute costs remain relatively low regardless of task complexity. It replaces older architectures that require fixed, high resource allocations.
A new model called JEPA-Anything works across physics, biology, and medicine. It predicts future states by splitting them into multiple partial parts.
Training - The massive reinforcement learning run on Nvidia GPUs
Beam's capabilities stem from a unique training phase involving reinforcement learning. The team ran this process on 10,500 Nvidia GB300 GPUs over a period of four weeks. This setup represents one of the largest training runs conducted by any open lab group.
The reinforcement learning phase focused heavily on reasoning and coding tasks. Performance metrics continued to improve throughout the entire duration of the run. There were no signs of plateauing even after exceeding 80 million rollouts.
Rollouts refer to the number of times the model attempts a task during training. The sheer volume of attempts allowed the system to refine its decision-making logic. Standard pretraining alone could not achieve the observed performance levels on these benchmarks.
Users can control how thoroughly the model reasons through specific problems. A tunable parameter lets operators choose between speed and accuracy trade-offs. This feature allows engineers to balance compute cost against output quality dynamically.
Measured by token generation counts, Beam works more efficiently than competitors. It achieves similar performance levels while generating fewer tokens overall. This efficiency extends across all three key benchmarks the company highlights. The training run demonstrated that massive scale does not always equal maximum efficiency.
Performance - Benchmark scores against GLM 5.2 and Qwen 3.8
Beam matches GLM 5.2 on several key industry benchmarks according to Reflection. These comparisons cover coding tasks, logical reasoning, and agent capabilities. The open model comes close to the much larger Qwen3.8-Max variant.
Stronger open models like Kimi K3 still beat Beam on raw performance metrics. The company acknowledges this gap but claims it is narrowing with future iterations. Reflection states it is already training a successor to close the remaining difference.
Across all three benchmarks, Beam delivers comparable scores while using far less compute. This combination of speed and accuracy makes it attractive for cost-sensitive projects. The data suggests that raw parameter count does not dictate final model quality.
Engineers should compare these scores against their own internal testing results. Context windows and specific task requirements may alter the relative performance. The company plans to publish a full technical report with detailed evaluation methods soon.
The benchmarks focus on practical applications rather than abstract reasoning tests. This emphasis helps businesses understand real-world value for their AI investments. Coding and agent tasks represent the primary use cases for this new model.
Capabilities - Emergent skills in web browsing and tool use
During training, Reflection observed what it calls emergent capabilities appearing in Beam. The model improved at web browsing even though no such tasks existed in its training mix. This phenomenon suggests complex behaviors can arise from simpler learning processes.
While running a mix of reasoning, software engineering, and terminal tasks, the company noticed this shift. With access to the web, the model independently learned to query other language models. It also pulled documents from external services without explicit programming for these actions.
Reflection shared several live demos showcasing these capabilities in action. One demo featured a live-updating New York City subway map visualization. Another included a small 3D game rendered through text-based interfaces. A third example showed a notebook for fine-tuning another AI model.
Beam is strictly text-only but can process content from other media formats. As long as images or audio are represented as text, the model handles them correctly. This limitation means users must convert non-text data into readable strings first.
The emergent skills demonstrate the model's ability to adapt to new environments. It does not require retraining when presented with novel tools or interfaces. This flexibility is valuable for agents needing to interact with dynamic systems.
Safety - A separate model handles alignment and hard rules
For safety and alignment, Reflection trained a second specialized model and merged it with Beam. The guidelines range from hard rules the model must never break to quality standards. These standards include factual accuracy and admitting uncertainty when facts are unknown.
The company also defined a direct, thorough, and proactive response style for all interactions. Safety test results will be published in a technical report alongside the model weights. Reflection plans to open-source the evaluation methods it developed for these tests.
Beam is still going through final safety testing before full public release. An early version is currently available to select users who want to try it first. The team aims to ensure the model adheres to strict ethical guidelines before wider adoption.
A separate model handles alignment and hard rules
For safety and alignment, Reflection trained a second specialized model and merged it with Beam. The guidelines range from hard rules the model must never break to quality standards. These standards include factual accuracy and admitting uncertainty when facts are unknown.
The company also defined a direct, thorough, and proactive response style for all interactions. Safety test results will be published in a technical report alongside the model weights. Reflection plans to open-source the evaluation methods it developed for these tests.
Beam is still going through final safety testing before full public release. An early version is currently available to select users who want to try it first. The team aims to ensure the model adheres to strict ethical guidelines before wider adoption.
Background - The founders and funding behind Reflection
Reflection was founded in 2024 by former Google Deepmind researchers Misha Laskin and Ioannis Antonoglou. Laskin previously led reward modeling for the Gemini project at Google. Antonoglou helped build AlphaGo, the system that defeated world champions in Go.
The startup launched in March 2025 with $130 million in seed funding. Their goal was to build superintelligence through autonomous coding systems. They envisioned language models learning to act as independently as AlphaGo plays games.
In summer 2025, the company released Asimov, an agent for analyzing large codebases. Reflection then raised $2 billion at an $8 billion valuation in October 2025. Nvidia joined the investors alongside other major tech companies and cloud providers.
The company has since positioned itself as a Western counterpart to Deepseek and Qwen. It releases model weights but keeps training data and pipelines proprietary. More recently, Reflection signed billion-dollar compute deals with SpaceX and Nebius.
These partnerships secure the massive GPU resources needed for future model iterations. The funding allows the team to continue pushing boundaries in reinforcement learning. Their vision remains focused on autonomous coding and general-purpose agents.
Release details - License, pricing, and availability timeline
Beam will ship under the Apache 2.0 open-source license later this month. This license permits free use, modification, and distribution of the model weights. Developers can integrate Beam into their own products without paying licensing fees.
The company states that technical documentation and developer guides will arrive simultaneously with the weights. Pricing remains a topic for individual enterprise negotiations rather than public listing. Early adopters may face different terms compared to later general releases.
An early version is available to select users for now while final testing continues. The full release includes all safety checks and evaluation reports from the team. Reflection encourages engineers to check their specific use cases against the guidelines before deployment.
The availability timeline suggests a phased rollout rather than an immediate global launch. This approach allows the company to monitor real-world performance and address issues quickly. Users should expect potential delays if critical bugs are found during testing.
What to do - How engineers can access and test the model
Engineers interested in Beam should start by reviewing the official developer documentation once it ships. The team will provide clear instructions on how to download the model weights locally. Testing on a small dataset is recommended before running large-scale production workloads.
Users can compare Beam's performance against existing models using standard evaluation frameworks. Reflection plans to open-source its evaluation methods, making benchmarking easier for everyone. This transparency helps the community verify claims about compute efficiency and reasoning quality.
Select users can access an early version today while waiting for the full release. These testers provide valuable feedback that shapes the final product's safety and performance. The company values input from engineers who will actually run these models in production.
Compare Beam's benchmarks against GLM 5.2 and Qwen 3.8 using your own internal metrics. Consider how the mixture-of-experts architecture fits your current infrastructure constraints. If cost is a primary concern, Beam offers a compelling alternative to larger competitors.
Check the safety guidelines carefully before integrating the model into sensitive applications. The separate alignment model ensures hard rules are followed consistently across all interactions. Monitoring logs should be enabled to track adherence to these safety protocols over time.
Reflection AI released Beam, a 501B parameter MoE model with 23B active parameters. It targets coding and agentic workloads but is not yet available for self-hosting.