As reported by The Logic, with details from Oak Lab.
The father of reinforcement learning just founded a lab whose first principle is everything today’s frontier AI conspicuously lacks — and the timing is not a coincidence.
What Happened
Richard Sutton — the researcher most widely called the father of reinforcement learning, co-winner of the 2024 Turing Award alongside Andrew Barto, and author of the field-defining essay The Bitter Lesson — has founded a new AI lab. Oak Lab (@oaklab_ai) is Canadian-incorporated and co-founded with his former student Khurram Javed. Sutton is departing John Carmack’s Dallas AGI startup Keen Technologies to build it, a move first reported by The Logic. The departure follows Google DeepMind’s closure of the Edmonton lab Sutton helped lead — the institutional home of some of the most consequential RL research of the past two decades.
Oak’s stated goal is architecturally unusual in the current market: agents that learn in real time without storing or replaying data, improving continually while using less processing power. The lab’s founding document — A Vision of Superintelligence from Experience — argues that intelligence is created through run-time experience, not pre-training on human-generated text, and that all components of a genuine intelligence system must learn continually. Current deep-learning methods, LLMs included, are characterized as too weak and too inefficient to reach real intelligence without fundamentally new ideas.
This is a direct architecture bet, not a product launch. Oak is an early-stage research lab with no product, no disclosed funding round, and no public timeline. The intellectual provocation is deliberate and specific — and it lands in a week that already delivered sharp empirical context for the argument Sutton is making.
The key insight: Sutton’s Bitter Lesson was widely read as a pro-scale manifesto. Oak Lab is the correction: the lesson was always about general methods and search and learning — not about who can buy the most GPUs. The lab is Sutton making that distinction impossible to ignore.
The Structural Read
Three forces converge in the Oak Lab founding, and they compound each other in ways the headline obscures.
1. The critique and the cure arrived the same week. Days before Sutton’s founding went public, a benchmark study — covered here in the AI Skyfall analysis — showed that frontier LLMs including GPT-5 and Gemini do not learn on the job. They run frozen heuristics. They cannot adapt when their environment changes between calls. The clinical term is “catastrophic forgetting at the deployment layer.” Oak’s founding principle — that all components of an intelligent system must learn continually from experience — is a direct, named solution to a demonstrated, measured deficiency. That is not a coincidence of timing; it is the market creating the conditions for the counter-argument to find air.
2. Algorithms over compute is a direct repricing event — if it works. The entire industry is spending as if the path to intelligence is more GPUs and more scale. The $1T+ capex supercycle is underwritten by a single assumption: that current deep-learning methods, applied at sufficient scale, will produce the next qualitative leap. Oak’s promise of learning on less processing power — not as a product feature but as a philosophical premise — challenges the investment thesis behind every hyperscaler data center announcement of the past eighteen months. If a better algorithm requires structurally less compute to achieve equivalent or superior learning, then a material portion of the current buildout is misallocated capital. Sutton is not saying “scale is useless.” He is saying the current algorithm is the wrong thing to scale.
3. The contrarian thesis is now a funding category. Oak joins a small but serious cluster of “neolabs” — SSI, Thinking Machines, Yann LeCun’s world-model push (examined in our LeCun / Meta world-models analysis), Carmack’s Keen Technologies, now Oak — that are attracting capital specifically because they reject the consensus architecture. The pattern across all of them is identical: the durable asset each lab is chasing is a learning loop the incumbents cannot buy off the shelf. OpenAI, Anthropic, and Google DeepMind have distribution, users, and revenue — but their core models are static weights. A lab that cracks continual, on-device, experience-driven learning owns something that cannot be replicated by purchasing more H100s. That is the learning-loop moat — and it is the structural prize that makes these bets fundable even without near-term revenue.
The War of Agents Architectures
“The frontier labs are building increasingly capable systems on a fixed-weight foundation. The neolabs are betting that the foundation itself is the constraint. Both can be right for different time horizons — but only one architecture produces an agent that gets smarter every time it runs.”
The honest hedge: continual reinforcement learning as a path to AGI is not a new idea. It is Sutton’s idea, pursued across decades of serious work by him and his collaborators, and it has not yet produced a system that competes with frontier LLMs on any benchmark that general users care about. The gap between “LLMs have a ceiling” and “RL-from-experience will break through it” is not closed by founding a lab. Oak is a research bet, not a result. The LLM scaling camp has users, revenue, infrastructure, and most of the field’s talent momentum. “LLMs are a dead end” is a real intellectual position held by serious people — Sutton, LeCun, and others — and it is a minority view. Calling any architecture “the path to superintelligence” remains speculative on all sides. This is worth watching closely; it is not worth treating as settled.
Three Implications
IMPLICATION 1 — Continual Learning Is the Missing Primitive
The Skyfall benchmark made the deficiency measurable; Oak is the first pure-play research lab organized entirely around closing it. If continual learning at inference time is the capability that unlocks the next generation of useful agents, then the lab that owns that primitive owns a structural position no amount of fine-tuning or retrieval-augmented generation can replicate. Watch for enterprise AI buyers — especially in dynamic, fast-changing environments like financial trading, robotics, and autonomous logistics — to treat this research thread as a strategic dependency within 24 months.
IMPLICATION 2 — The Capex Assumption Has a Named Critic With Standing
Until now, the loudest skeptics of the compute supercycle were economists and short-sellers. Sutton is something different: a 2024 Turing Award winner with a specific, testable alternative hypothesis and a lab organized to prove it. If Oak publishes results — even early, partial results — showing that a continual-learning RL agent beats a same-generation LLM on adaptive tasks with less compute, that is a data point that moves infrastructure investment narratives. It does not need to be AGI to matter; it needs to be real and reproducible.
IMPLICATION 3 — The Neolab Era Is a Structural Bifurcation, Not a Blip
SSI, Thinking Machines, Oak, and the LeCun world-model cluster are not failed OpenAI applications. They are a deliberate architectural fork. Capital is now bifurcating: frontier-scale investments for the LLM incumbents, and smaller but serious bets on the labs building the architecture that the incumbents cannot pivot to without dismantling their core product. The War of Agents Architectures and The Agentic AI Stack are the right frames for tracking how this fork resolves over the next three to five years.









