Reflection AI Unveils Beam: A 501B Open-Weight Model Built for Efficient Reasoning
AI News

Reflection AI Unveils Beam: A 501B Open-Weight Model Built for Efficient Reasoning

5 min
10/6/2026
Reflection AIBeamopen-weight modelMixture-of-Experts

Reflection AI Enters the Open-Weight Arena with Beam

On October 5, 2026, Reflection AI announced Beam, its first open-weight model, marking a significant entry into the competitive open-source AI landscape. Beam is a sparse Mixture-of-Experts (MoE) model with 501 billion total parameters but only 23 billion active per token. This architecture is designed to deliver high capability while keeping inference costs manageable.

The model was built with a dual focus on pretraining and reinforcement learning (RL). Reflection pretrained Beam on 23.8 trillion curated tokens from the web and proprietary licensed datasets. The company then conducted a massive high-compute RL run, generating over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks. This combination aims to produce a model that is both knowledgeable and adept at complex reasoning and agentic tasks.

Capabilities and Benchmark Performance

Beam's development prioritized coding and agentic workloads, and the benchmark results reflect this focus. On SWEBench Verified, Beam scored 80.9, surpassing models like Inkling (77.6) and Nemotron 3 Ultra (70.7). On Terminal Bench v2.1, it achieved 80.1, competitive with GLM 5.2 (81.0) and approaching Qwen 3.8-Max (86.6).

In reasoning tasks, Beam posted strong numbers: 97.8 on AIME 2026, 90.5 on GPQA Diamond, and 36.2 on Humanity's Last Exam (without tools). These scores place it in the upper tier of open-weight models, though it trails frontier models like Kimi K3 on raw capability. Reflection acknowledges this, positioning Beam's advantage as inference-time efficiency rather than absolute performance.

The Efficiency Edge: 3-4x Less Compute

A key selling point for Beam is its efficiency. Reflection claims Beam achieves scores comparable to GLM 5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. This efficiency is even more pronounced when compared to larger models like Qwen 3.8-Max, which require significantly more compute per token.

This efficiency is partly due to the MoE architecture—only 23B parameters are active per token—and partly due to training innovations. Reflection implemented a controllable length penalty during RL that rewards successful solutions while discouraging unnecessary tokens. This results in a model that can adapt its reasoning effort based on the task, offering a "reasoning effort parameter" for users to balance speed and depth.

continue reading below...

Scaling Reinforcement Learning to New Heights

Reflection's RL infrastructure is a major technical achievement. The company deployed 10.5K NVIDIA GB300 GPUs for four weeks, generating over 100 million rollouts with a maximum context length of 256K tokens. They used approximately 1.3 billion sandboxes and sourced one million high-quality coding, agentic, and STEM environments.

To maintain stability at this scale, Reflection developed new algorithms to handle policy staleness in asynchronous policy gradients. Even with one-day staleness—107 weight versions behind the current policy—the training remained numerically stable. The infrastructure also supported an average of 110K concurrent rollouts, with fast model updates (median 12 seconds) and resilience to inference failures (71 incidents handled without terminating training).

Pretraining and Architecture Innovations

Beam's pretraining was completed in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. Reflection built nearly the entire infrastructure stack in-house, including a topology-aware scheduler and a silent data corruption detection system. The run achieved 92.3% goodput towards the end, a testament to the robustness of their systems.

Architecturally, Beam incorporates interleaved local and global attention, fine-grained routed experts, and multiple load-balancing techniques. The result is near-uniform expert utilization (busiest expert at 1.04x load), which ensures all experts are effectively used during reasoning. The residual stream is carefully managed with depth-based scaling and FP32 accumulation to maintain healthy signal propagation across all 52 layers.

Safety and Alignment Approach

Reflection used a multi-teacher on-policy distillation (MOPD) method to merge capabilities from a large-scale RL teacher and a dedicated safety teacher. The safety principles are organized into three tiers: rules not to break, qualities to consistently satisfy, and default interaction style.

The safety training used deliberative alignment techniques, incorporating the safety policy directly into the model's reasoning. The dataset was built adversarially, with each round generating prompts that elicited harmful or over-refusing behavior and folding successful attacks back into the training mix. Reflection plans to open-source the safety evaluations they developed internally.

Availability and Roadmap

Beam is currently in final red-teaming, with early access available through a waitlist on Reflection's platform. The full weights, technical report, model card, and developer artifacts are scheduled for release later in October 2026 under an Apache 2.0 license.

Reflection plans to launch Beam with a broad ecosystem of distribution partners and integrations with open-source libraries. This is the first model in a planned series, with Reflection already training its successor. The goal is to bring the open frontier closer to the frontier of intelligence with each release.

Beam represents a significant step for Western open-weight AI, offering a competitive alternative to models from Chinese labs like GLM and Qwen. Its efficiency and strong coding/agentic performance make it a compelling option for enterprises and developers looking for a powerful yet cost-effective workhorse model.