Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Reflection AI launches Beam open-weight model with 23B active parameters - OpenSmartRoute
Reflection AI launches Beam open-weight model with 23B active parameters
Reflection AI released Beam, a 501B parameter MoE model with 23B active parameters. It targets coding and agentic workloads but is not yet available for self-hosting.
Key points
Beam uses 3 to 4x less inference compute than larger open models on reasoning benchmarks.
It scores 80.9 on SWE-bench Verified, beating Nemotron 3 Ultra at 70.7.
The model features 501B total parameters and 23B active parameters per token.
Apache 2.0 weights are scheduled for release later this month.
Why it matters: Engineers can deploy a high-performance coding model that costs significantly less to run than current leaders.
By OpenSmartRoute editorial · written through the router by writer-small
From MarkTechPost - “Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads”
Reflection AI launched Beam, its first open-weight model. The company released details about a new large language model designed for coding and agent tasks. This model uses a specific architecture to balance speed and intelligence. It aims to compete with existing giants while using less compute power. Engineers can now see the technical specifications behind this release.
Beam is a Mixture-of-Experts model built by Reflection AI. The team states it targets enterprise workloads involving complex coding and reasoning. It does not yet allow users to host the software on their own servers. Early adopters must sign up for a waitlist on the Reflection platform. The company is currently conducting final red-teaming tests before full release.
Announcement - Reflection AI introduces Beam as its first open-weight model
Reflection AI has officially announced the existence of Beam. This marks a significant shift in their product roadmap and public offerings. They have chosen to make the weights available under permissive licenses. The goal is to advance the Western open-weight frontier for developers. Many competitors focus on closed systems that hide their internal logic. Beam breaks this pattern by sharing its core training data and architecture details.
The model directly competes with larger open models like GLM 5.2. Reflection positions Beam as a more efficient alternative for specific use cases. They acknowledge that Kimi K3 currently leads in raw capability metrics. The pitch focuses on efficiency at the moment of inference time. Users who care about cost per token will find this distinction important.
Architecture details - Sparse MoE structure with interleaved attention and routed experts
Beam uses a sparse Mixture-of-Experts design to process input tokens. Only a small fraction of parameters activate for every single token generated. This approach keeps the total parameter count high while maintaining fast inference speeds. The architecture interleaves local attention mechanisms with global attention layers. Fine-grained routed experts handle different parts of the user request dynamically.
Mistral launched Mistral Large 4, nicknamed Le Chonk. It is a 1 trillion-parameter model available for free.
Load balancing relies on auxiliary-loss-free balancing techniques from previous models. Cosine decay of expert-bias updates helps distribute work evenly across the network. The busiest expert reached just 1.04 times the average load during training. This stability prevents bottlenecks that often plague large language model systems. Across all 52 layers, residual norms stayed bounded using specific scaling methods. SandwichNorm and attention gating helped manage the depth-based scaling effectively. FP32 residual accumulation ensured numerical precision throughout the computation graph.
Training specifics - Token counts, GPU clusters, pretraining duration, and RL scale
The model was pretrained on 23.8 trillion tokens gathered from various sources. These included web data, public datasets, and proprietary licensed collections. The curation process removed about 95% of raw internet tokens to improve quality. Roughly 1.8 trillion high-quality tokens remained after filtering out noise. Conventional filters would have dropped many of these valuable training examples.
Pretraining finished in under four weeks using a massive GPU cluster. The system utilized 6,144 NVIDIA GB300 NVL72 GPUs for the entire duration. Goodput reached 92.3% near the end of the training run. Nine semi-automatic rewinds occurred to correct minor errors during the process. Midtraining extended the effective context window to 1 million tokens. This allowed the model to learn from longer sequences of text and code.
Reinforcement Learning serves as the central scaling axis for Beam's final tuning. The run used 10.5K NVIDIA GB300 GPUs over a four-week period. It generated over 100 million rollouts to refine the model's behavior. Maximum rollout context reached 256K tokens for complex scenario testing. Training and grading consumed about 1.3 billion sandboxes across environments. These covered nearly one million coding, agentic, and STEM domains.
Benchmark results - Scores on SWE-bench, Terminal Bench, and comparison to rivals
On the SWE-bench Verified benchmark, Beam scores 80.9 points. This compares favorably to Nemotron 3 Ultra which scores 70.7 points. On Terminal Bench v2.1, Beam achieves a score of 80.1 points. GLM 5.2 scores close to this at 81.0 points on the same test. DeepSeek V4.1 Flash and Kimi K3 lead with higher scores there.
These numbers come from Reflection's official table of results. The team sources rival scores from Artificial Analysis and DataCurve. Beam is the smallest model here by total parameter count. Its 23B active count sits below GLM-5.2, Nemotron 3 Ultra, and Kimi K3. Despite being smaller, it matches performance on critical coding benchmarks. This suggests high efficiency in how it utilizes its active parameters.
Why it matters - Cost efficiency and performance trade-offs for enterprise workloads
Enterprise teams care deeply about the cost per token for their applications. Beam uses 3 to 4 times less inference compute on reasoning benchmarks. This reduction translates directly to lower operational expenses for large deployments. Companies running agents can match effort to task difficulty more precisely. Lower settings favor short answers while higher settings allow longer reasoning chains.
The model offers a controllable length penalty for solving tasks with fewer tokens. Browsing skills improved without explicit browsing tasks in the RL mix. This suggests strong transfer across different agentic domains and tools. Safety training used deliberative alignment to ensure responsible behavior. Reflection trained a separate safety teacher from the pretrained checkpoint initially. They merged that teacher with the RL teacher using multi-teacher distillation.
Safety evaluation results will appear in the upcoming technical report. The team reports no plateau as RL compute increased significantly. Infrastructure numbers show the system sustained 110K concurrent rollouts on average. New weights reached the inference fleet in a median of about 12 seconds. Seventy-one inference incidents were handled without stopping training operations.
How it compares - What existed before, what this changes and what stays the same
Beam replaces the need for massive closed models in coding tasks. It competes with GLM 5.2 on reasoning benchmarks directly. Beam uses significantly less compute than its rivals during inference. Existing open models like Llama 3 still dominate general chat applications. Kimi K3 leads on raw capability scores across most major tests.
Beam stays smaller than the top tier of current market leaders. Its total parameter count is 501 billion, which is lower than many competitors. The active parameters per token sit at 23 billion for efficiency reasons. This design choice targets specific enterprise coding and agentic workloads. It does not replace general-purpose chat models for casual users yet.
The architecture differs from standard dense transformer designs. Beam interleaves local attention with global attention mechanisms effectively. Fine-grained routed experts handle different tasks within the same layer. Load balancing relies on auxiliary-loss-free balancing techniques from DeepSeek-V3. This approach keeps expert utilization near average levels consistently.
Pretraining data curation removed about 95% of raw internet tokens. The team kept roughly 1.8 trillion high-quality tokens for training. Conventional filters would have dropped these specific high-value tokens. This selective process improved the model's reasoning capabilities significantly. It avoids the noise found in uncurated web scraping datasets.
Reflection trained a separate safety teacher from the pretrained checkpoint initially. They merged that teacher with the RL teacher using multi-teacher distillation. Safety training used deliberative alignment methods to ensure responsible behavior. The team reports no plateau as RL compute increased significantly over time. This stability allows for continuous improvement without performance drops.
Questions this leaves open - What the source does not say and how a reader can check it
Readers need to verify the actual inference latency numbers in practice. The article mentions reduced compute but not specific millisecond timings. Engineers must test Beam against their own local hardware environments first. Self-hosting is currently unavailable for external teams through official channels.
The licensing terms for commercial use require closer legal review soon. Apache 2.0 and MIT licenses are scheduled for later this month. Kimi K3's custom license adds attribution requirements for very large products. Beam follows a similar pattern with specific legal considerations for commercial use. Teams must confirm these details before integrating into production systems.
Safety evaluation results will appear in the upcoming technical report soon. The current announcement lacks detailed breakdowns of safety failure modes. Independent auditors should review the full report for comprehensive alignment checks. Red-teaming is still ongoing within the Reflection team internally. Early access runs through a waitlist on the Reflection platform only.
The 1M effective context window needs real-world validation for long documents. Midtraining extended effective context to 1M tokens during specific phases. Users should test token limits with their longest expected input files. Context compression strategies may differ from standard dense model behaviors.
Benchmark scores come from Reflection's table which sources rival scores. Artificial Analysis and DataCurve provide the underlying data for these comparisons. Independent verification of SWE-bench Verified and Terminal Bench v2.1 results is possible. DeepSeek V4.1 Flash and Kimi K3 lead there with higher scores currently.
The 10.5K NVIDIA GB300 GPUs used for RL training are not publicly available yet. Infrastructure numbers show the system sustained 110K concurrent rollouts on average. New weights reached the inference fleet in a median of about 12 seconds. This speed depends entirely on specific hardware availability and network conditions.
Seventy-one inference incidents were handled without stopping training operations during development. Teams need to understand how their own systems handle similar edge cases. Monitoring tools must track expert load balancing under heavy traffic loads. The busiest expert reached just 1.04x average load at the end of pretraining.
The reasoning effort parameter allows teams to match effort to task difficulty. Lower settings favor short answers while higher settings allow longer reasoning chains. Users should calibrate these settings based on their specific compute budgets. Cost savings come from avoiding unnecessary computation on simple tasks.
Reflection positions Beam as advancing the Western open-weight frontier specifically. The research team is candid about the gap compared to Kimi K3. Kimi K3 stays ahead on raw capability according to the team's own assessment. Efficiency at inference time remains Beam's primary pitch and selling point.
Readers should check the Technical details page for deeper architectural insights soon. The Early Access portal will open once red-teaming concludes officially. Teams should monitor the release schedule for updates on licensing terms regularly. Asif Razzaq is the CEO of Marktechpost AI Media Inc. He provides context on the broader landscape of open model development.
The platform covers machine learning news that is technically sound and accessible to many. Over 2 million monthly views illustrate its popularity among audiences globally. Beyond Domain-Specific World Models use a single recipe for seven different fields currently. Meet Together Link offers a free CLI that runs open models inside various environments.
Yandex Introduces Sona as a single generative recommender that replaces entire recommendation cascades. The Story of Qwen traces Alibaba's AI models from 7B to 2.4T parameters over time. These examples show the rapid pace of innovation in the current AI landscape. Engineers must stay updated on emerging architectures and training methodologies constantly.
Beam represents a shift toward specialized efficiency rather than pure scale dominance. Managers can assess the total cost of ownership before committing to adoption fully. The focus remains on practical performance metrics rather than raw scale alone currently. Companies running agents can match effort to task difficulty more precisely now.
What to do - Waitlist access, license details, and monitoring the release schedule
Users interested in early access should join the waitlist on the Reflection platform. The company has not yet enabled self-hosting for external teams. Apache 2.0 and MIT are standard permissive licenses available soon. Kimi K3's custom license adds attribution requirements for very large products. Beam follows a similar pattern with specific legal considerations for commercial use.
Check out the Technical details page for deeper architectural insights. The Early Access portal will open once red-teaming concludes. Teams should monitor the release schedule for updates on licensing terms. Asif Razzaq is the CEO of Marktechpost AI Media Inc. He provides context on the broader landscape of open model development. His platform covers machine learning news that is technically sound and accessible.
The announcement highlights a clear path forward for developers seeking efficiency. Reflection positions Beam as a viable alternative to current market leaders. The focus remains on practical performance metrics rather than raw scale alone. Engineers can evaluate the model against their specific workload requirements now. Managers can assess the total cost of ownership before committing to adoption.