Reflection AI Debuts Beam, a 501B Open-Weight Agent Model
The Nvidia-backed startup introduces a 501B MoE model with 23B active parameters, which Reflection AI claims matches GLM-5.2 on advanced reasoning benchmarks at 3-4x less inference compute, while trailing other Chinese models like Kimi K3 on raw capability.

Key takeaways
- Beam features 501 billion total parameters but activates only 23 billion per token, reducing inference compute requirements.
- Trained on 23.8 trillion tokens using thousands of Nvidia GB300 GPUs across pretraining and reinforcement learning phases.
- Weights will be released under an Apache 2.0 license later in October 2026 following final red-teaming and safety evaluations.
Nvidia-backed AI startup Reflection AI announced Beam on October 5, 2026, marking its debut in the open-source artificial intelligence ecosystem. The system is a sparse Mixture-of-Experts (MoE) foundation model totaling 501 billion parameters, engineered specifically for software development, logical reasoning, and autonomous agent tasks. The model uses an active parameter routing mechanism to keep runtime resource demands manageable.
According to an announcement covered by The Next Web, Beam activates only 23 billion parameters per token out of its 501 billion total and features a 1 million token context window.

Architecture and Pretraining Specs
Reflection AI built Beam as a text-only, 52-layer sparse Mixture-of-Experts architecture. Technical details compiled by MarkTechPost highlight several design choices to maintain stability across large GPU clusters, including interleaved local and global attention, SandwichNorm, depth-based residual scaling, attention gating, and FP32 residual accumulation.
The model incorporates auxiliary-loss-free load balancing with cosine decay of expert-bias updates, keeping the average load on the busiest expert to 1.04 times the mean. Beam's base pretraining consumed 23.8 trillion tokens curated from web and commercial datasets, executed on 6,144 Nvidia GB300 NVL72 GPUs in under four weeks with a 92.3% goodput rate.
Data filtering removed roughly 95% of raw internet tokens while retaining approximately 1.8 trillion high-quality tokens—including 87% of curated web code tokens—that standard heuristic filters typically discard, according to Unite.AI. A subsequent midtraining phase expanded the effective context length to 1 million tokens.
Massive Reinforcement Learning Pipeline
Following base pretraining, Reflection AI deployed an extensive reinforcement learning campaign spanning four weeks on 10,500 Nvidia GB300 GPUs. The company utilized approximately 1.3 billion sandboxes across nearly 1 million coding, STEM, and agentic environments.
As reported by MarkTechPost, the reinforcement learning infrastructure generated more than 100 million rollouts with context sizes reaching up to 256,000 tokens. The startup, which raised funding at a $25 billion valuation, previously signed compute agreements including a $6.3 billion deal with SpaceX and a $1 billion deal with Nebius.
According to Reflection AI, during training, Beam demonstrated spontaneous cross-domain transfer: despite the absence of dedicated browsing tasks in its RL mixture, the model improved its web-navigation capabilities and independently learned to invoke optical character recognition APIs and query external models when given web access.

Benchmark Comparisons and Inference Efficiency
Reflection AI positioned Beam as a high-efficiency alternative to larger open-weight models, particularly those originating from Chinese research labs that have recently led open benchmarks.
According to evaluation tables published by Reflection and detailed by MarkTechPost, Beam scored 80.1 on Terminal Bench v2.1, matching Z.ai's 753B parameter GLM 5.2 (81.0) and outperforming Nvidia's Nemotron 3 Ultra (56.4), though trailing Moonshot AI's 2.8-trillion parameter Kimi K3 (88.3) and DeepSeek V4.1 Flash (90.6).
On SWE-bench Verified, Reflection AI reported that Beam achieved 80.9, topping Nemotron 3 Ultra's 70.7 and Thinking Machines Lab's Inkling score of 77.6, though these scores have not been independently verified. The company noted that Beam matches GLM 5.2 reasoning scores while utilizing three to four times less inference compute, calculated based on active parameter counts and average generated tokens.
Safety Alignment and Release Schedule
To manage safety without degrading core reasoning performance, Reflection AI trained a separate safety and alignment teacher checkpoint using deliberative alignment principles. This model was subsequently merged into the primary policy using multi-teacher on-policy distillation.
Chief executive Misha Laskin stated that two government evaluation groups—the US Center for Advancing Innovation and Standards for Super Intelligence and the UK AI Safety Institute—are assessing Beam ahead of public availability. Beam is currently accessible to select users through a waitlist while undergoing final red-teaming.
Reflection AI plans to publish the model weights, evaluation tools, and technical documentation under a permissive Apache 2.0 license later in October 2026. The founders confirmed that training for their next-generation architecture is already underway.
Frequently asked questions
What is Reflection AI Beam?
Beam is an open-weight mixture-of-experts artificial intelligence model developed by Reflection AI, featuring 501 billion total parameters and 23 billion active parameters per token.
When will Beam's weights be available to download?
Reflection AI plans to release the model weights, model card, and fine-tuning stack under an Apache 2.0 license later in October 2026.
What hardware was used to train Beam?
Beam was pretrained on 6,144 Nvidia GB300 NVL72 GPUs in under four weeks, followed by a four-week reinforcement learning run on 10,500 GB300 GPUs.
Sources
- Nvidia-backed Reflection unveils Beam, its first open-weight modelTNW | Artificial-intelligence · Oct 5, 2026
- Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloadsmarktechpost.com · Oct 5, 2026
- Reflection AI debuts open-source Beam model with 501B parametersSiliconANGLE · Oct 6, 2026
- Reflection AI Unveils Beam, a 501B-Parameter Open-Weight ModelUnite.AI · Oct 5, 2026
- This new AI model could help America close a technological gap with ChinaMorningstar, Inc. · Oct 5, 2026
How this story was made: the newsroom picked it up from Google News, Bluesky and siliconangle.com, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (33 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 6, 2026 at 01:10 UTC


