Reflection AI spent more compute on reinforcement learning for its Beam model than on pretraining, co-founder and CEO Misha Laskin said on the No Priors podcast in an episode published on 9 October 2026. Pretraining used 6,000 GB300s, while reinforcement learning was “a little over 10,000 GB300s for four weeks,” he said.
Laskin described Beam as “a 500 billion parameter model, total 23B active.” He said the pretraining run took a few weeks, and that with infrastructure efficiencies it can now be done in about 12 days, maybe less.
Business Pill · SCALING LAWS
A short explainer of scaling laws: why AI labs keep building bigger. It teaches the general idea only and says nothing about any company in this story.
The key insight: As we read it, Laskin’s figures put the bigger compute bill after pretraining. On his numbers, reinforcement learning took more chips for longer than a pretraining run at the 12-day pace he described, about four times the GB300-weeks on our arithmetic.
The Compute Split
“So actually, there were more flops spent on reinforcement learning,” Laskin said, calling the faster pretraining “a coupling of scientific and infrastructure efficiencies.”
On our arithmetic from his figures, the reinforcement learning run comes to about 40,000 GB300-weeks (10,000 chips for four weeks), against roughly 10,300 GB300-weeks for a pretraining run at the 12-day pace he described. He gave the original pretraining run only as a few weeks, so that run cannot be computed exactly.

The Cost of Catching Up
Answering a question about the resources needed to catch up to the frontier, Laskin said it would have been the order of a hundred million dollars a year to 18 months ago, and “probably order billion, so billions of dollars, single-digit billions” now or six months ago.
“Going into next year, I think order 10,” he said, adding that “for every generation of model, there’s a 4x multiplier in compute, roughly.”

Open Models on the Gateways
Laskin said that about six months ago, traffic on gateways such as OpenRouter or Vercel was majority closed models, and that it has “flipped almost exactly from 70-30 closed open to 70-30 open closed now.” He said he expects that to accelerate.
He compared the direction to operating systems, saying that 95% plus of servers and computers in the world run on an open source operating system like Linux.
The Structural Read
Post-training is now the larger run. Laskin said more flops went to reinforcement learning than to pretraining for Beam.
Efficiency compresses one stage. He said a pretraining run that took a few weeks can now be done in about 12 days, maybe less.
The entry price keeps rising. He described the cost of catching up moving from hundreds of millions of dollars to single-digit billions, with roughly 4x compute per model generation.
Misha Laskin, Reflection AI, No Priors, 9 October 2026
“So actually, there were more flops spent on reinforcement learning.”
Three Implications
RL ON 10,000-PLUS GB300S Laskin said reinforcement learning ran on a little over 10,000 GB300s for four weeks.
GATEWAYS TILTING OPEN He said traffic on gateways such as OpenRouter or Vercel flipped to about 70-30 open versus closed.
ROUGHLY 4X PER GENERATION Laskin put the compute multiplier per model generation at roughly 4x.
The Business Engineer Lens
This story maps onto the Business Engineer framework The Open vs Closed Meta-Framework.
The framework’s starting point: “Close the SCARCE layer (keep proprietary) + Open the ABUNDANT layer (commoditize).”
As we read it, Laskin’s account places the scarce layer in compute and post-training, the reinforcement learning run being the bigger of the two, while the episode presents Beam as an open model. On his description of gateway traffic, open models are where usage is moving.
What Is Not Established
We read transcripts of three segments of the episode; we could not retrieve the full episode captions, so the rest of the conversation was not read. Every figure here is Laskin’s on-record statement, not audited data, and we did not check the gateway shares against OpenRouter or Vercel data. We did not contact Reflection.
The Bottom Line
Laskin says Reflection spent more compute on reinforcement learning for Beam, a little over 10,000 GB300s for four weeks, than on pretraining on 6,000 GB300s. He puts the cost of catching up to the frontier at single-digit billions of dollars now, up from hundreds of millions, with “order 10” next year.
94,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. We read transcripts of three segments of the No Priors episode of 9 October 2026 in which Misha Laskin speaks; the full captions were not retrievable. The GPU-week comparison is our arithmetic. We did not contact Reflection. Nothing here is a forecast, and nothing here is financial or investment advice.









