Reflection previews 501-billion-parameter Beam model ahead of open-weight release
Beam activates 23 billion parameters at a time and targets coding and agentic work, but its weights, model card and technical report are still promised for later in October.

The story
Reflection AI has unveiled Beam, its first large language model, presenting a 501-billion-parameter mixture-of-experts system designed for coding, reasoning and tool-using agents. The announcement puts a heavily financed US laboratory into the contest to build capable models whose weights can be downloaded and adapted rather than accessed only through a vendor-controlled service. For now, however, Beam is a preview: Reflection is offering early access to selected users while final red-teaming and evaluations continue.
Beam contains 501 billion parameters in total but activates 23 billion for each token. That sparse architecture is intended to deliver the capacity of a very large network without paying the full computational cost on every inference step. Reflection says the text-only model has an effective context length of one million tokens and was pretrained on 23.8 trillion tokens drawn from the web, public material and proprietary licensed datasets. The company says it removed about 95% of raw internet tokens during parsing, deduplication and quality filtering, although it has not yet published a complete training-data inventory.
The reported training scale is substantial. Reflection says Beam's pretraining finished in under four weeks on 6,144 Nvidia GB300 NVL72 GPUs. A separate reinforcement-learning campaign used 10,500 GB300 GPUs for four weeks, generated more than 100 million rollouts and involved roughly 1.3 billion sandbox executions. The laboratory says its asynchronous system sustained an average of 110,000 concurrent rollouts and trained on a pool of nearly one million coding, agentic and STEM environments.
Reflection's central performance claim is efficiency. Its published results place Beam near Z.ai's GLM-5.2 and closer to Alibaba's Qwen3.8-Max on selected coding and agentic tests, while acknowledging that Kimi K3 remains ahead on raw capability. The company estimates that Beam matches GLM-5.2 on advanced reasoning while using three to four times less inference compute. Those comparisons are company-reported rather than independently verified, and Reflection's own methodology describes them as approximate because it excludes prompt processing, context-dependent attention and serving overhead.
That qualification matters. The launch post includes a broad benchmark table, but no downloadable weights, technical report, model card or complete safety-evaluation results were available at announcement. TechCrunch similarly noted that the performance claims had not been independently verified. Until external researchers can run the same checkpoints under controlled conditions, the figures should be treated as evidence of Reflection's positioning—not a settled ranking of open models.
Reflection says it will release Beam's weights later in October under the Apache 2.0 license, together with documentation and a stack for running, evaluating and fine-tuning the model. If delivered as described, that package would give companies and public institutions more control over deployment, adaptation and data locality than a closed application programming interface. Open weights are not identical to fully open-source development, however: meaningful transparency also depends on the detail of the training recipe, data disclosures, evaluation code and reproducibility artifacts.
The geopolitical context is as important as the architecture. Reuters reported that Reflection is explicitly positioning Beam against lower-cost Chinese models such as DeepSeek and Kimi, while the company says it wants to advance a Western open-weight frontier. Chinese laboratories have repeatedly shown that capable, adaptable models can reshape pricing and deployment choices worldwide. A competitive US-developed alternative could broaden the supplier base for enterprises and governments that want to operate models in their own infrastructure.
The model's size still sets a high practical bar. Activating 23 billion parameters reduces computation compared with running all 501 billion at once, but hosting the full checkpoint, moving expert weights and serving long contexts will require substantial memory, networking and systems engineering. Reflection has not yet published minimum hardware requirements, measured serving throughput, deployment pricing or independently reproducible cost comparisons. Those details will determine whether Beam is broadly usable or mainly attractive to hyperscalers, specialist clouds and sovereign computing programs.
INNOVOX analysis: Beam is consequential less because it wins a single benchmark than because it combines frontier-scale US capital, current-generation Nvidia hardware and a stated Apache-licensed release plan. The sparse design and high-compute reinforcement-learning campaign show how the competitive frontier is shifting from model size alone toward the efficiency of training, reasoning and agent execution. Yet this remains a promise-backed preview. Its significance will rise sharply only if the public release is complete enough for outsiders to inspect, reproduce and deploy.
What to watch next is the actual October release. The decisive evidence will be the exact license, checkpoint availability, tokenizer and inference code, model card, safety results, data-provenance disclosures and a technical report detailed enough to audit the benchmark methodology. Independent tests should measure coding reliability, tool-use failure modes, long-context accuracy, memory requirements, tokens per second and total serving cost. Any difference between the preview claims and the released artifacts will matter more than the launch-day leaderboard.
INNOVOX analysis
Beam could widen the Western supply of adaptable high-end models, but launch-day benchmarks cannot establish its position. The key test is whether Reflection ships complete, inspectable artifacts that external teams can reproduce and operate economically.
What to watch
Watch for the checkpoint and exact license, model card, technical report, data disclosures, independent benchmark replications, minimum hardware requirements, measured serving throughput and total deployment cost.
