Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Reflection launches Beam open model with 501B parameters - OpenSmartRoute
Reflection launches Beam open model with 501B parameters
Reflection released Beam, a 501 billion parameter text-only model for coding and science. Apache 2.0 weights are available this month after training on 23.8 trillion tokens.
Key points
Beam uses 501B total parameters with 23B active in a Mixture of Experts design.
The team trained on 23.8T pretraining tokens including hundreds of millions of PDFs.
Reflection paid $150M/month for Colossus compute and $1B for Nebius storage.
Beam scores 80.9 on SWE-bench Verified with claims of four weeks of RL training.
Why it matters: Engineers get a new US-trained open model option while managers track the $1B+ compute cost behind it.
By OpenSmartRoute editorial · written through the router by writer-small
From Latent Space - “[AINews] Reflection Beam - 501B-A23B American Open Model”
Reflection launches Beam as a functional open model
Reflection finally released its Beam model after more than a year of silence. The team had hinted at coding goals but kept their launch plans secret. Many observers assumed they would never ship a working model like this. They compared themselves to Inkling, Nemotron, and GLM 5.2 in the past. However, models like GLM 5.3 and DeepSeek V4.1 generally lead the field now. This US-trained model fills a specific gap for developers. The broader ecosystem did not slow down big releases from other labs. Reflection has now announced its arrival as a functional open model provider.
Beam is a text-only model designed for coding, agency, and science work. It uses 501 billion total parameters with 23 billion active parameters in its mixture of experts. Full weights are available under the Apache 2.0 license this month. The announcement came from Laskin regarding the release details. Training data includes over 23.8 trillion pretraining tokens. Part of this massive dataset comes from an OCR pipeline processing hundreds of millions of PDFs. Data lead confirmed the inclusion of these documents in the training mix.
The model was trained from scratch without fine-tuning on existing proprietary datasets. This approach aims to create a unique capability for American developers. The team plans to release tech reports and open-source integrations soon. Polozov mentioned these upcoming OSS integrations as part of their roadmap. The goal is to provide a robust alternative to closed models currently dominating the market. Engineers can expect more transparency in how the model performs on specific tasks.
Beam architecture and training scale details
The architecture uses a stable reinforcement learning approach called OPD for optimization. Damos described this run as involving 10,000 GB300 chips for the training infrastructure. The system executed more than 100 million rollouts across approximately one million distinct tasks. This scale allowed the model to learn complex coding patterns effectively. Elie Bakouch estimated the model achieves roughly 12% BF16 MFU during pretraining. MFU measures how efficiently compute resources are utilized for training workloads.
The attention mechanism features a 3:1 interleaved global and sliding-window design. This structure helps balance long-range dependencies with memory efficiency. The team reports better held-out code perplexity compared to DeepSeek V4.1. Bakouch noted this metric as a key indicator of model quality in coding scenarios. Perplexity measures how surprised the model is by the text it generates during testing.
A new model called JEPA-Anything works across physics, biology, and medicine. It predicts future states by splitting them into multiple partial parts.
Training costs and infrastructure choices
Axios reported that Reflection pays $150 million per month for Colossus compute resources. They also secured a separate $1 billion deal with Nebius cloud services. These massive expenditures support their ability to train such large models from scratch. Other unnamed US labs are expected to ship open models this month as well. Curran noted the broader trend of American labs entering the open model space.
Independent analysis suggests Beam will be among the most token-efficient open models. Artificial Analysis expects high intelligence per token spent on inference. Teortaxes calls Beam an iso-FLOP replication of DeepSeek V3 in terms of compute usage. Iso-FLOP means the model uses the same amount of floating-point operations as its counterpart.
Compute comparison shows Beam runs roughly 1.3 billion RL sandboxes over a four-week period. Up to 170,000 sandboxes can run simultaneously during this training phase. Sandboxes provide isolated environments for testing agent behaviors safely. This parallel execution is crucial for scaling reinforcement learning experiments.
Benchmark results and competitive positioning
Claimed results include an 80.9 score on SWE-bench Verified, a standard coding benchmark. The model also achieves three to four times the inference efficiency of GLM 5.2. Inference efficiency measures how many tokens the model processes per unit of time or cost. Four weeks of pretraining and RL ran on approximately 10,500 GB300 chips.
Observers place Beam around the performance level of GLM-5.2 on certain benchmarks. However, it trails DeepSeek V4.1 Flash on some specific metrics according to critics. Nathan Lambert groups it with Nvidia and Thinking Machines as strong US releases. These models still trail Chinese counterparts in overall benchmark scores.
Aleph Alpha Kolibri offers a 78 billion total parameter model for German and English tasks. It reports 96.9% accuracy on AIME 2025 and 84.3% on GPQA Diamond. Self-reported scores are high, but the dataset remains unreleased to the public. Agentic evaluations sit well below Qwen in comparative tests.
Decision models like Command Code's Agr skip text generation entirely. It returns typed values with per-option probabilities for tool calls and routing. SemiAnalysis explains that TypeSafe's Jev uses the same no-decode approach. These models mainly displace frontier models in router roles rather than general chat.
Smaller releases include Upstage's Solar Mini 4 with a 35 billion parameter count. It offers a 512K context window and is free on Nous Portal for two weeks. Eleven v4 Turbo tops AA's Provider Voice TTS arena at half the price of v4. These smaller models target specific niches like voice synthesis or limited tasks.
Why it matters for engineers and managers
Beam addresses the lack of US-trained open models in the current market. Developers have been eagerly waiting for more options beyond Chinese or European providers. The Apache 2.0 license allows commercial use without restrictive copyright clauses. Managers can evaluate the model against competitors like GLM 5.3 and Qwen 3.8 Max.
Engineers benefit from the specific focus on coding and scientific workloads. General models often struggle with complex code generation tasks requiring precision. The high token efficiency reduces operational costs for large-scale deployments. Cost savings become significant when processing millions of tokens daily.
Safety concerns remain a priority for any new open model release. The team avoided fine-tuning on proprietary data to maintain transparency. Independent analysis checks the model's alignment and reasoning capabilities before public launch. Engineers should verify performance on their specific codebases before adoption.
What to do with the new model and data
Readers can check the full weights under Apache 2.0 this month for immediate access. The Latent Space website hosts the official announcement and technical details. Compare Beam's SWE-bench Verified score against other coding models in your stack. Look at inference efficiency metrics if cost is a primary concern for your team.
Try integrating the model into your existing agent workflows via promised OSS integrations. Polozov indicated these integrations will simplify deployment for many users. Monitor performance on held-out code perplexity to gauge generalization quality. Use independent benchmarks like SWE-bench to validate claims made by the team.
Explore the 23.8 trillion token dataset if you have access to research materials. Understanding the data composition helps explain the model's strengths and weaknesses. Check OCR pipeline details to see how much unstructured text was included. This information is vital for understanding potential biases in the training process.
Reflection launches Beam as a functional open model
Reflection announced Beam today. It is a new text-only model. The team calls it a functional neolab. They trained it from scratch using only American data. Full weights arrive under Apache 2.0 this month. This license permits commercial use freely. No restrictive copyright clauses limit your business. The launch ends a long period of stealth. Reflection stayed quiet longer than Thinking Machines. Many users waited for a US option. Chinese models dominated the recent open model space. GLM 5.3 and Qwen 3.8 Max lead that space. Beam aims to fill the gap here. It targets coding, agentic work, and science. The team compares itself to Inkling and Nemotron. They acknowledge these competitors are strong. Still, the US market needs more choices. Latent Space confirmed the news on their site. AINews checked 12 subreddits for updates. They found no further Discords or Twitters. The announcement date is October 3, 2026.
Beam architecture and training scale details
Beam uses a specific MoE design. It has 501 billion total parameters. Only 23 billion parameters are active at once. This architecture balances speed and capacity. Elie Bakouch analyzed the structure closely. He notes a 3:1 attention ratio. Global attention mixes with sliding-window attention. This pattern helps hold code perplexity low. The team cites 23.8 trillion pretraining tokens. Part of this data comes from OCR. Hundreds of millions of PDFs feed the pipeline. A stable RL/OPD run followed training. They used 10,000 GB300 chips for this phase. More than 100 million rollouts crossed ~1 million tasks. Damos described these rollout statistics. The team claims four weeks of pretraining time. RL training also took four weeks. Both phases ran on roughly 10,500 GB300s. Teortaxes calls this an iso-FLOP replication. He compares it to DeepSeek V3 directly.
Compute costs and infrastructure choices
Axios reported the compute costs ahead of time. Reflection pays $150 million per month for Colossus. They also signed a $1 billion deal with Nebius. Other unnamed US labs ship open models this month. Curran noted these parallel launches. One analyst claims Anthropic spends 42% on subscriptions. These subscriptions earn only about 10% of revenue. The chart shows this spending ratio clearly. OpenAI faces capacity squeezes during high demand. New $200 sign-ups were paused recently. Usage limits effectively halved across plans. GPT-6.1 Sol positions itself as an efficient alternative. Theo describes a reversal in coding model preference. July and September saw different trends. Codex lead Tibo pledged daily improvements or resets. He promised this for 28 days straight. Day 1 speedup hit default speeds for GPT-6 Astra. Speed rose from ~30 to ~50 TPS. This change covers all subscription surfaces. Sign in with ChatGPT partners saw the boost. OpenCode, Pi, Amp, and Devin included here. Friction remains with banked Codex resets. They expire without timezone adjustment. The always-on dots agent limits Pro plans. It requires $100+ Pro subscriptions to run.
Benchmark results and competitive positioning
Claimed results show 80.9 on SWE-bench Verified. This score beats many coding benchmarks. Beam offers 3–4x the inference efficiency of GLM 5.2. Four weeks of training took ~10,500 GB300s. Polozov promised tech reports and OSS integrations. Artificial Analysis expects high token efficiency. They view it as one of the most efficient open models. MFU estimates only ~12% BF16 MFU in pretraining. Bakouch reads this as a significant constraint. Held-out code perplexity beats DSv4 according to analysis. Observers place Beam around GLM-5.2 level. Some benchmarks put it below DSv4 Flash. Nathan Lambert groups it with Nvidia releases. He calls them strong US releases that trail Chinese counterparts. Aleph Alpha Kolibri reports 96.9% AIME 2025 scores. It has 78 billion total parameters. The dataset remains unreleased to the public. Agentic evals sit well below Qwen according to Jitsev. Reka Rho-1 trained on 320 H100s in ~3 months. Command Code's Agr skips text generation entirely. It returns typed values with probabilities for tool calls. SemiAnalysis explains TypeSafe's Jev uses the same approach. These models displace frontier models mainly in router roles. Upstage's Solar Mini 4 offers a 35 billion parameter count. It provides a 512K context window. Eleven v4 Turbo tops AA's Provider Voice TTS arena.
Why it matters for engineers and managers
Beam addresses the lack of US-trained open models. Developers waited eagerly for more options. The Apache 2.0 license allows commercial use freely. Managers can evaluate the model against competitors like GLM 5.3. Engineers benefit from the specific focus on coding workloads. General models often struggle with complex code generation tasks. High token efficiency reduces operational costs for large-scale deployments. Cost savings become significant when processing millions of tokens daily. Safety concerns remain a priority for any new open model release. The team avoided fine-tuning on proprietary data to maintain transparency. Independent analysis checks the model's alignment and reasoning capabilities before launch. Engineers should verify performance on their specific codebases before adoption.
What to do with the new model and data
Readers can check the full weights under Apache 2.0 this month. The Latent Space website hosts the official announcement and technical details. Compare Beam's SWE-bench Verified score against other coding models in your stack. Look at inference efficiency metrics if cost is a primary concern for your team. Try integrating the model into your existing agent workflows via promised OSS integrations. Polozov indicated these integrations will simplify deployment for many users. Monitor performance on held-out code perplexity to gauge generalization quality. Use independent benchmarks like SWE-bench to validate claims made by the team. Explore the 23.8 trillion token dataset if you have access to research materials. Understanding the data composition helps explain the model's strengths and weaknesses. Check OCR pipeline details to see how much unstructured text was included. This information is vital for understanding potential biases in the training process.
Reflection AI released Beam, a 501B parameter MoE model with 23B active parameters. It targets coding and agentic workloads but is not yet available for self-hosting.