Reflection AI Unveils Beam, a 501B Open-Weight Reasoning Model
The Nvidia-backed startup is aiming its 501-billion-parameter mixture-of-experts model at enterprise coding, agentic workflows, and sovereign AI deployments.

Key takeaways
- Reflection AI introduced Beam, a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active parameters per token.
- The model was pretrained on 23.8 trillion tokens and refined through a four-week reinforcement learning run across 10,500 Nvidia GB300 GPUs.
- Reflection claims Beam matches leading models like GLM 5.2 on advanced reasoning while using three to four times less inference compute.
- Model weights, technical documentation, and fine-tuning tools are scheduled for release later in October 2026 under an Apache 2.0 license.
Reflection AI introduced Beam on October 5, 2026, marking its debut open-weight model engineered for coding, multi-step reasoning, and agentic workflows. Built as a sparse mixture-of-experts architecture, Beam contains 501 billion total parameters but activates only 23 billion parameters per token. The Brooklyn-based startup positions the system as an efficient Western alternative to major closed systems and leading open-weight releases, promising high-tier reasoning capabilities at reduced compute overhead.
The launch represents the first in a planned series of releases from Reflection AI, founded in 2024 by former Google DeepMind researchers and backed by Nvidia, Sequoia Capital, and Lightspeed Venture Partners, according to TechCrunch. The company plans to release the model weights, evaluation harnesses, and full technical documentation under an Apache 2.0 license later in October 2026 following final red-teaming.

Architecture and Pretraining Efficiency
Beam relies on a sparse mixture-of-experts design that interleaves local and global attention across 52 layers, as outlined by MarkTechPost. Its foundation was pretrained on 23.8 trillion tokens drawn from web corpora, public repositories, and proprietary licensed datasets. The training pipeline filtered out roughly 95% of raw internet tokens, while Reflection AI said conventional filtering techniques would have missed roughly 1.8 trillion high-quality tokens it retained, including 87% of its curated web-code tokens.
Reflection completed the pretraining run in under four weeks on an infrastructure cluster of 6,144 Nvidia GB300 NVL72 GPUs. To maintain stability across the network, the engineering team implemented depth-based residual scaling, SandwichNorm, elementwise attention gating, and FP32 residual accumulation. Load balancing was adapted from DeepSeek-V3's auxiliary-loss-free method with the addition of cosine decay on expert-bias updates, resulting in near-uniform routing where the busiest expert load averaged 1.04 times the mean by the end of training.
Pretraining goodput reached 92.3% of wall-clock time despite nine semi-automatic rewinds prompted by gradient-norm spikes or suspected silent data corruption. A subsequent midtraining phase expanded Beam's effective context window to one million tokens.
High-Compute Reinforcement Learning and Emergent Skills
To develop task execution and multi-step reasoning capabilities, Reflection deployed a reinforcement learning campaign across 10,500 Nvidia GB300 GPUs over a four-week period, generating more than 100 million rollouts, according to Unite.AI. The team constructed an environment pool of nearly one million synthetic, proprietary, and open-source coding, agentic, and STEM tasks, executing roughly 1.3 billion training and grading sandboxes.

The training platform sustained an average of 110,000 concurrent rollouts using fully asynchronous policy gradients. Tokens were tagged with their generating policy version, allowing learning to remain numerically stable even when training on samples up to 107 versions (approximately 24 hours) behind the current model weights. The system handled 71 inference incidents without interrupting training runs, and updated weights propagated across inference nodes in a median of 12 seconds.
During reinforcement learning, Reflection observed that Beam developed agentic capabilities outside its explicit training mixture. Although no browsing tasks were included in the RL mix, the model improved its web browsing performance and autonomously learned to query external models and employ text-recognition tools to parse documents. To manage resource consumption, Reflection incorporated a controllable length penalty, enabling users to adjust a reasoning effort parameter that balances token generation against task difficulty.
Benchmark Performance and Inference Efficiency
According to benchmarks reported by The Next Web and Reflection's technical post, Beam scored 80.9 on SWE-bench Verified and 80.1 on Terminal Bench v2.1. By comparison, Nvidia's Nemotron 3 Ultra scored 70.7 on SWE-bench Verified and 56.4 on Terminal Bench v2.1, while Thinking Machines Lab's Inkling registered 77.6 on SWE-bench Verified and 63.8 on Terminal Bench v2.1.
Against leading international open models, Beam sits close to Z.ai's GLM-5.2 (81.0 on Terminal Bench v2.1) and approaches Alibaba's Qwen 3.8-Max, though Moonshot AI's Kimi K3 (88.3) and DeepSeek V4.1 Flash (90.6) maintain higher raw benchmark scores. Reflection highlighted that Beam achieves comparable scores to GLM-5.2 while requiring three to four times less inference compute on advanced reasoning benchmarks. Independent evaluation firm Artificial Analysis noted that the 23-billion active parameter configuration positions Beam among the most token-efficient open models evaluated for its intelligence tier, as reported by 24/7 Wall St..

Alignment, Enterprise Strategy, and Availability
For safety and alignment, Reflection trained a dedicated teacher model from the base checkpoint using deliberative alignment and multi-turn adversarial red-teaming, distilling it into the final policy alongside the reinforcement learning teacher. Co-founder and CEO Misha Laskin noted that both the US Center for Advancing Innovation and Standards for Super Intelligence and the UK AI Safety Institute are reviewing the model prior to public distribution.
Beam forms the core of Reflection's commercial "AI factory" architecture, designed to let enterprises and sovereign entities train and host proprietary models locally. Reflection has already established infrastructure partnerships, including compute commitments totaling more than $7 billion with Nebius and SpaceX, a deployment partnership with Dell AI Factory, consortium participation in the US Department of Energy's Genesis Mission, and a 250-megawatt sovereign AI cloud initiative with South Korea's Shinsegae Group, according to TechCrunch and Unite.AI.
While self-hosting weights will arrive later in October 2026 under the Apache 2.0 license, developer access is currently available via an early-access waitlist on Reflection's platform.
Frequently asked questions
What is Reflection AI Beam?
Beam is a 501-billion-parameter sparse mixture-of-experts open-weight AI model developed by Reflection AI, activating 23 billion parameters per token for coding, reasoning, and agentic workflows.
When will Beam's weights and code be publicly released?
Reflection AI plans to release the model weights, technical documentation, and fine-tuning tools later in October 2026 under the Apache 2.0 license, following completion of red-teaming evaluations.
How does Beam achieve lower inference costs?
By using a sparse MoE architecture with only 23 billion active parameters per token and applying length penalties during reinforcement learning, Reflection AI estimates Beam uses 3x to 4x less generation compute than GLM-5.2 on advanced reasoning benchmarks based on its approximate calculations.
What hardware was used to train Beam?
Beam was pretrained on 6,144 Nvidia GB300 NVL72 GPUs in under four weeks, followed by a four-week reinforcement learning run on 10,500 Nvidia GB300 GPUs.
Sources
- Nvidia-backed Reflection unveils Beam, its first open-weight modelTNW | Artificial-intelligence · Oct 5, 2026
- Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloadsmarktechpost.com · Oct 5, 2026
- Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute costTechCrunch · Oct 5, 2026
- Reflection AI Unveils Beam, a 501B-Parameter Open-Weight ModelUnite.AI · Oct 5, 2026
- Reflection's Beam model draws an independent token-efficiency verdict24/7 Wall St. · Oct 5, 2026
- OpenAI and Anthropic Have a New Threat to Worry About—and It Isn't Chinagizmodo.com · Oct 5, 2026
- Introducing Beam: Reflection’s 501B open-weight model — ReflectionReflection
How this story was made: the newsroom picked it up from Google News, Techmeme and Bluesky, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (49 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 6, 2026 at 00:39 UTC


