ModelsArchitectures & capability
Reflection AI announces Beam, a 501B-parameter open-weight model it says rivals GLM-5.2 on less compute
Reflection AI says its 501B-parameter open-weight model, Beam, matches China's GLM-5.2 on hard reasoning tests using 3–4× less inference compute. The claim is self-reported, its own table shows rivals ahead on several benchmarks, and weights remain unreleased.

Beam is a text-only sparse mixture-of-experts model with 501B total parameters and 23B active per token. Reflection pitches it for reasoning, coding and agentic work. For now, access is through a waitlist. Reflection says Apache 2.0 weights, a technical report, a model card and safety results will follow later in October 2026, after red-teaming is finished. [2] [1]
The headline claim is that Beam is on par with GLM-5.2 at 3–4× lower inference compute. That figure comes from Reflection alone. Reflection's own table shows rivals scoring higher on some tests, including SWE-Bench Pro v2-Hard and AIME 2026. Its coding win over Inkling compares a text-only model with a multimodal one. No independent evaluation exists yet. [1] [2] [3] [7]
Beam supports Reflection's business plan: building custom 'AI factories' for customers, with Shinsegae Group in South Korea as the first sovereign test. The company is Nvidia-backed and raised about $2B in 2025 at an $8B valuation. It has also signed more than $7B in deals with SpaceX and Nebius for GB300 compute through 2029. [1] [3] [5]
only vendor benchmarks and specs exist so far. Implication: if Apache 2.0 weights ship as planned, Beam could give buyers wary of Chinese-origin models a permissive US alternative, though hosting still requires memory for all 501B parameters.
Executive brief
Reflection AI has announced Beam, a 501B-parameter open-weight model. The company says it matches China's GLM-5.2 on hard reasoning benchmarks while using 3–4× less inference compute. No one outside the company has tested that claim yet. Reflection's own comparison table also shows Chinese models ahead on several benchmarks. The weights are not public. Reflection says it will release them under Apache 2.0 later in October 2026, and early access is through a waitlist (Reflection blog; TechCrunch). Reflection is two years old, has raised about $4.7B, and Nvidia is one of its backers.
What changed and event timeline
$2B raise for an open frontier lab
Reflection raised about $2B at an $8B valuation from investors including Nvidia, Lightspeed and Sequoia, and said it would build open frontier models as a Western counterpart to DeepSeek ().
Summer 2026: Compute locked up through 2029
Reflection signed deals worth more than $7B in total with SpaceX and Nebius for access to Nvidia GB300 chips (). SiliconANGLE puts the SpaceX portion at $6.3B ().
Later Oct 2026 (planned): Weights and technical report
Reflection says the weights, technical report, model card and safety benchmarks will come out after final red-teaming ().
Axios reports a launch is close
Axios said a launch was imminent and that hedge funds and trading firms were interested in Reflection's "AI factory" concept ().
Beam announced
Reflection's blog post gave the specs, training compute and benchmark claims, and opened a waitlist. Reflection did not answer TechCrunch's follow-up questions in time for publication (;).
Capabilities and access
- Model: Beam, a text-only sparse mixture-of-experts (MoE) model with 501B total parameters, of which 23B are active per token.
- Context window: TechCrunch reports 1M tokens. Reflection's blog says 256K, extended to 1M during midtraining (TechCrunch).
- Access: waitlist and early access for now. Reflection plans Apache 2.0 weights, distribution through hyperscalers and neoclouds, and integrations with open-source libraries.
Read the full section
- Model: Beam, a text-only sparse mixture-of-experts (MoE) model with 501B total parameters, of which 23B are active per token. Reflection pitches it as a "workhorse" for reasoning, coding and agentic tasks (Reflection blog).
- Context window: TechCrunch reports 1M tokens. Reflection's blog says 256K, extended to 1M during midtraining (TechCrunch).
- Access: waitlist and early access for now. Reflection plans Apache 2.0 weights, distribution through hyperscalers and neoclouds, and integrations with open-source libraries.
Technical analysis for researchers and developers
- Architecture: attention alternates between local and global layers.
- Pretraining: 23.8T tokens of web and licensed data, run on 6,144 GB300 GPUs for about 4 weeks.
- Reinforcement learning: about 10.5K GB300 GPUs for 4 weeks, with more than 100M rollouts across about 1M coding, agentic and STEM environments, plus 1.3B sandbox evaluations.
Read the full section
The following is from Reflection's blog (Reflection blog):
- Architecture: attention alternates between local and global layers. The model uses fine-grained routed experts, a "controlled residual stream" and load balancing to keep expert use stable.
- Pretraining: 23.8T tokens of web and licensed data, run on 6,144 GB300 GPUs for about 4 weeks.
- Reinforcement learning: about 10.5K GB300 GPUs for 4 weeks, with more than 100M rollouts across about 1M coding, agentic and STEM environments, plus 1.3B sandbox evaluations.
- Post-training: separate SFT and RL pipelines, merged through multi-teacher on-policy distillation.
- What the efficiency estimate leaves out: prompt prefill, context-dependent attention cost and serving overhead. Real deployment costs may therefore differ from the 3–4× figure.
- Reproducibility: none yet. Neither the weights nor the technical report are public.
Claims and evidence
No independent evaluation of Beam exists yet (Dataist).
Read the full section
| Claim | Status |
| On par with GLM-5.2 on reasoning, using 3–4× less inference compute | Vendor-reported; not verified independently (TechCrunch) |
| Beats Thinking Machines' Inkling on four coding tests | Vendor-reported. Inkling is multimodal and Beam is text-only, so the comparison is not like-for-like (TechCrunch) |
| Comes close to Qwen 3.8-Max (2T+ parameters) but trails Claude Fable 5.1 | Vendor-reported (SiliconANGLE) |
| Rivals score higher in Reflection's own table (e.g., SWE-Bench Pro v2-Hard 77.2 vs. DeepSeek V4.1 Flash 84.3; AIME 2026 97.8 vs. GLM 5.2 99.2) | Vendor's table (Reflection blog) |
No independent evaluation of Beam exists yet (Dataist).
Context and prior work
- Founders: Reflection was started in 2024 by former DeepMind researchers Misha Laskin (who led reward modeling for Gemini) and Ioannis Antonoglou (co-creator of AlphaGo).
- Rival models by size: Beam is smaller than GLM-5.2 (about 744B total, 40B active).
- Competitors: Chinese open models (DeepSeek, Qwen, Z.ai, Kimi) and Western ones (Mistral, Meta, Cohere, and Thinking Machines' Inkling, released in July) (TechCrunch).
Read the full section
- Founders: Reflection was started in 2024 by former DeepMind researchers Misha Laskin (who led reward modeling for Gemini) and Ioannis Antonoglou (co-creator of AlphaGo). The company first worked on coding agents (The AI Insider).
- Rival models by size: Beam is smaller than GLM-5.2 (about 744B total, 40B active).
- Competitors: Chinese open models (DeepSeek, Qwen, Z.ai, Kimi) and Western ones (Mistral, Meta, Cohere, and Thinking Machines' Inkling, released in July) (TechCrunch).
Limitations, safety and contested findings
- Benchmarks are self-reported. Reflection's own table shows Beam trailing named rivals on several tests, so "matches Chinese models" only holds for some benchmarks (Reflection blog).
- The efficiency figure is partial. It leaves out prefill and serving overhead.
- Safety results are not out. Red-teaming is still underway, and the safety results are due with the technical report.
Read the full section
- Benchmarks are self-reported. Reflection's own table shows Beam trailing named rivals on several tests, so "matches Chinese models" only holds for some benchmarks (Reflection blog).
- The efficiency figure is partial. It leaves out prefill and serving overhead.
- Safety results are not out. Red-teaming is still underway, and the safety results are due with the technical report.
- Sources disagree on context length: 256K native (blog) vs. 1M (TechCrunch).
- AI factory is unproven. There is no evidence yet that the concept works at scale (Dataist).
Business and practitioner implications
- Licensing: if the release happens as promised, Apache 2.0 makes Beam a permissively licensed US option for buyers who are wary of Chinese-origin weights.
- Serving cost: with 23B active parameters, per-token compute is lower, but hosting still needs memory for all 501B parameters.
- Business model: Reflection plans to sell "AI factories," meaning custom local models trained on a customer's own data.
Read the full section
- Licensing: if the release happens as promised, Apache 2.0 makes Beam a permissively licensed US option for buyers who are wary of Chinese-origin weights.
- Serving cost: with 23B active parameters, per-token compute is lower, but hosting still needs memory for all 501B parameters.
- Business model: Reflection plans to sell "AI factories," meaning custom local models trained on a customer's own data. Its first sovereign test is with Shinsegae Group in South Korea, and the concept fits Nvidia's push for more GPU demand (TechCrunch).
- Practical step: run your own evaluations once the weights ship before committing.
Sources
The source trail.
Sources (7)
TechCrunch: Reflection debuts Beam
techcrunch.comReflection: Introducing Beam
reflection.aiSiliconANGLE: Reflection AI debuts Beam
siliconangle.comCoverage notes no independent evaluation of Reflection's Beam yet
Related coverage; assess separately
cryptobriefing.com