- Reflection AI says its Beam open-weight model matches GLM-5.2 on advanced reasoning benchmarks while using 3–4x less inference compute. The claim is not independently verified.
- For buyers, the useful measure is cost per solved task, not benchmark rank. Reflection plans to release the weights under Apache 2.0 later this month.
- Once the weights ship, run your own tasks at several reasoning-effort settings and compare cost per correct result.
Beam is sold on running cost, not top score
A fuel-economy rating tells a fleet manager more than a top-speed figure. Reflection AI's pitch for its new Beam model works the same way. The company does not claim the most capable open model. It claims a cheaper one to run. The question for buyers is what each correct answer costs.
Reflection announced Beam on October 5. It is the startup's first open-weight model. Open weight means the trained model files can be downloaded and run on your own hardware. Reflection says it will release them under an Apache 2.0 license later this month. Until then, Beam is in final red-teaming and available only through an early-access waitlist.
What Reflection claims, and what is unverified
Reflection says Beam is competitive with larger open models, including GLM 5.2, on coding and agentic tasks. Agentic means the model plans and uses tools over many steps. On advanced reasoning benchmarks, the company reports scores comparable to GLM-5.2's. It says Beam needs three to four times less computing at answer time to get there. That answer-time computing is called inference.
Reflection also says Kimi K3 remains ahead on raw capability. Its pitch is efficiency, not the top score.
TechCrunch reports that these performance claims have not been independently verified. Reflection did not respond to TechCrunch's requests for more information in time for publication. Broad independent testing has to wait for the public weights.
How a large model can be cheap to run
Beam is built from many specialist sub-networks, called experts. For each word fragment, or token, the model routes work to only a few of them. This design is known as sparse Mixture-of-Experts. Beam has 501 billion parameters in total but uses 23 billion at a time, under 5% of the whole.
TechCrunch gives GLM 5.2 as roughly 744 billion total and 40 billion active. Fewer active parameters means less computing per token. That is one plausible source of the saving. Reflection's second lever is shorter reasoning. It trained Beam with a length penalty that pays out for successful solutions and pushes against padding.
Buyers get a dial for this: a reasoning effort setting. Lower settings give shorter answers. Higher settings allow longer reasoning on hard tasks. The cost of a task therefore depends on how you configure the model, not only on its price list.
The training behind it
Reflection trained Beam with reinforcement learning, in which the model practises tasks and is scored on the results. The company says the run used 10.5K NVIDIA GB300 GPUs over four weeks and produced more than 100 million rollouts, or full attempts at a task. It drew on nearly one million task environments. Reflection believes few open labs have run reinforcement learning at this scale. That is the company's own assessment.
The scale of that run is a reminder that a model cheap to use can still be very costly to build. Reflection has not said what the training cost.
Reflection also reports that Beam got better at browsing although no browsing tasks were in its training mix. Given web access, it started on its own to run searches, ask other language models questions and call OCR tools to read documents. For site owners, this is a reminder that capable agents are becoming another kind of visitor. Reflection reports the behavior from its own testing.
Who is funding the bet
TechCrunch reports Reflection has raised roughly $4.7 billion, per PitchBook. Over the summer it also signed two compute contracts, with SpaceX and Nebius, that together exceed $7 billion. They reserve access to NVIDIA's GB300 chips into 2029. The company is aiming Beam at enterprises and sovereign nations. Its pitch is "AI factories": customised local systems trained on an institution's own data.
Questions to put to your team
First, wait for the weights and the technical report. Ask whether your team can run Beam on your own tasks, not Reflection's benchmarks.
Second, measure cost per correct result at several reasoning-effort settings. A cheaper token that needs more retries is not cheaper.
Third, ask for the safety evaluations. Reflection says it will publish its results in the technical report and open-source the evaluations it used internally. Check that these arrive before any production use.
Fourth, if you consider the "AI factory" idea, ask who carries the hardware, hosting and tuning costs. Reflection has described the concept but has not given pricing in the sources reviewed.
An efficiency claim is a hypothesis until someone outside the lab runs it. Open weights make that test possible. Run it before you sign.
Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Reflection AI.





