Reflection's Beam is a 501B open-weight model built for agents
Reflection has introduced Beam, a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active, claiming frontier-level coding and agentic results at a fraction of the inference cost.

Reflection has introduced Beam, its first open-weight model. Beam is a sparse mixture-of-experts system with 501 billion total parameters, of which 23 billion are active per token, and it is aimed squarely at coding, reasoning and agentic workloads. The company published the details on 5 October, with the weights, a technical report, a model card and the developer artefacts promised later this month and early access open in the meantime.
Pretrained on 23.8 trillion tokens
Beam was pretrained on 23.8 trillion tokens gathered from the web and from proprietary licensed datasets, which Reflection says matches or beats comparable open base models of the same size. The reinforcement-learning stage is where the bulk of the effort went: the company reports more than 100 million rollouts across four weeks on 10,500 Nvidia GB300 GPUs, a maximum context length of 256,000 tokens, and roughly 1.3 billion sandboxes used for training and grading. Reflection says it sourced a million coding, agentic and STEM environments to sustain the run, and claims it as one of the largest reinforcement-learning campaigns run by any open lab.
Where Beam sits on the benchmarks
Reflection's own table puts Beam at 80.9 on SWE-bench Verified, 80.1 on Terminal-Bench 2.1, 97.8 on AIME 2026 and 36.2 on Humanity's Last Exam without tools. The company describes Beam as competitive with GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, while conceding that frontier open models such as Kimi K3 remain ahead on raw capability. The argument it makes instead is efficiency: comparable reasoning scores to GLM-5.2 for three to four times less inference compute, and a wider gap again against models in the two-trillion-parameter class.
A crowded open-weight field
Beam arrives into a market where open weights are no longer a novelty. Chinese labs have set the pace for most of the past two years, and Western labs have spent 2026 trying to close the gap on capability while arguing they can win on cost per token. Reflection is one of a small group of vendors pitching sovereign, self-hosted deployment to enterprises that will not send code to a hosted API, and a model that is efficient at inference time is a more useful thing to sell to those customers than one that simply tops a leaderboard. There is no public pricing and no confirmed licence yet; both will matter as much as the benchmark numbers when the weights actually land.
Our opinion
The interesting number in Beam is not 501 billion, which is a parameter count doing marketing work, but the 23 billion that are active per token. Sparse models buy capability with memory and efficiency with routing, and Reflection is betting that buyers will pay for the second while the leaderboard rewards the first. That is a sensible read of the enterprise market, where the bill arrives every month and nobody has to know which model answered. It is also a bet that can be checked: publishing the weights later this month turns every claim in that benchmark table into something anyone with a pair of GPUs can attempt to reproduce. If the efficiency advantage holds up outside Reflection's own harness, Beam is a genuine shift in what an open model can cost to run. If it does not, the 501 billion figure will be the only part anyone remembers.