Checked fact 116811 Oct 2026Models and releases
Laskin said: "reinforcement learning was a little over 10,000 GB300s for four weeks" (12:47).
The exact words it rests on
12:47 efficiencies Um, reinforcement learning / 12:50 was / 12:52 a little over 10,000 GB300s for four / 12:54 weeks. So, actually there are more flops / 12:55 spent on reinforcement learning.
What the source said when we opened it, on 11 Oct 2026.
The source
Beam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
Checked
Checked by the notis newsroom on , against the source above.
In the story
Reflection AI's Laskin says open models will take most token demand 11 Oct 2026
Said aloud in a video: the quote, word for word.
Cite this fact
Anyone may quote this address. It does not change; if we correct the story, this page says so.