Oak Lab and Richard Sutton’s Bet Against the Scaling Paradigm

As reported by The Logic, with details from Oak Lab.

The father of reinforcement learning just founded a lab whose first principle is everything today’s frontier AI conspicuously lacks — and the timing is not a coincidence.

Oak Lab — The Signal Numbers

2024

Turing Award — Sutton & Barto (RL field-defining work)

2

Co-founders: Richard Sutton + Khurram Javed (former student)

0

Stored or replayed data — Oak’s stated architectural constraint

$1T+

Hyperscaler capex supercycle Oak’s algorithm thesis runs against

What Happened

Richard Sutton — the researcher most widely called the father of reinforcement learning, co-winner of the 2024 Turing Award alongside Andrew Barto, and author of the field-defining essay The Bitter Lesson — has founded a new AI lab. Oak Lab (@oaklab_ai) is Canadian-incorporated and co-founded with his former student Khurram Javed. Sutton is departing John Carmack’s Dallas AGI startup Keen Technologies to build it, a move first reported by The Logic. The departure follows Google DeepMind’s closure of the Edmonton lab Sutton helped lead — the institutional home of some of the most consequential RL research of the past two decades.

Oak’s stated goal is architecturally unusual in the current market: agents that learn in real time without storing or replaying data, improving continually while using less processing power. The lab’s founding document — A Vision of Superintelligence from Experience — argues that intelligence is created through run-time experience, not pre-training on human-generated text, and that all components of a genuine intelligence system must learn continually. Current deep-learning methods, LLMs included, are characterized as too weak and too inefficient to reach real intelligence without fundamentally new ideas.

This is a direct architecture bet, not a product launch. Oak is an early-stage research lab with no product, no disclosed funding round, and no public timeline. The intellectual provocation is deliberate and specific — and it lands in a week that already delivered sharp empirical context for the argument Sutton is making.

Sutton’s Intellectual Through-Line

2019 — The Bitter Lesson

Sutton argues that general methods leveraging computation always win over human-knowledge-encoded approaches. The essay becomes one of the most cited pieces in AI strategy — and is later misread as pure endorsement of scale.

2021 — Reward Is Enough

Sutton and colleagues publish the hypothesis that reward maximization — not pre-programmed world knowledge — is sufficient to produce all capabilities associated with intelligence. A direct shot across the architecture debate.

2025 — Era of Experience (with David Silver)

Sutton and Silver argue AI must move beyond data generated by humans to data generated by an agent’s own experience. Frames the LLM paradigm as a ceiling, not a frontier.

July 2026 — Oak Lab Founded

Sutton leaves Keen Technologies, incorporates Oak in Canada with Javed. Architecture: all components learn continually from experience, no data replay, less compute. The theory becomes a startup.

The key insight: Sutton’s Bitter Lesson was widely read as a pro-scale manifesto. Oak Lab is the correction: the lesson was always about general methods and search and learning — not about who can buy the most GPUs. The lab is Sutton making that distinction impossible to ignore.

The Structural Read

Three forces converge in the Oak Lab founding, and they compound each other in ways the headline obscures.

1. The critique and the cure arrived the same week. Days before Sutton’s founding went public, a benchmark study — covered here in the AI Skyfall analysis — showed that frontier LLMs including GPT-5 and Gemini do not learn on the job. They run frozen heuristics. They cannot adapt when their environment changes between calls. The clinical term is “catastrophic forgetting at the deployment layer.” Oak’s founding principle — that all components of an intelligent system must learn continually from experience — is a direct, named solution to a demonstrated, measured deficiency. That is not a coincidence of timing; it is the market creating the conditions for the counter-argument to find air.

2. Algorithms over compute is a direct repricing event — if it works. The entire industry is spending as if the path to intelligence is more GPUs and more scale. The $1T+ capex supercycle is underwritten by a single assumption: that current deep-learning methods, applied at sufficient scale, will produce the next qualitative leap. Oak’s promise of learning on less processing power — not as a product feature but as a philosophical premise — challenges the investment thesis behind every hyperscaler data center announcement of the past eighteen months. If a better algorithm requires structurally less compute to achieve equivalent or superior learning, then a material portion of the current buildout is misallocated capital. Sutton is not saying “scale is useless.” He is saying the current algorithm is the wrong thing to scale.

3. The contrarian thesis is now a funding category. Oak joins a small but serious cluster of “neolabs” — SSI, Thinking Machines, Yann LeCun’s world-model push (examined in our LeCun / Meta world-models analysis), Carmack’s Keen Technologies, now Oak — that are attracting capital specifically because they reject the consensus architecture. The pattern across all of them is identical: the durable asset each lab is chasing is a learning loop the incumbents cannot buy off the shelf. OpenAI, Anthropic, and Google DeepMind have distribution, users, and revenue — but their core models are static weights. A lab that cracks continual, on-device, experience-driven learning owns something that cannot be replicated by purchasing more H100s. That is the learning-loop moat — and it is the structural prize that makes these bets fundable even without near-term revenue.

The War of Agents Architectures

“The frontier labs are building increasingly capable systems on a fixed-weight foundation. The neolabs are betting that the foundation itself is the constraint. Both can be right for different time horizons — but only one architecture produces an agent that gets smarter every time it runs.”

The honest hedge: continual reinforcement learning as a path to AGI is not a new idea. It is Sutton’s idea, pursued across decades of serious work by him and his collaborators, and it has not yet produced a system that competes with frontier LLMs on any benchmark that general users care about. The gap between “LLMs have a ceiling” and “RL-from-experience will break through it” is not closed by founding a lab. Oak is a research bet, not a result. The LLM scaling camp has users, revenue, infrastructure, and most of the field’s talent momentum. “LLMs are a dead end” is a real intellectual position held by serious people — Sutton, LeCun, and others — and it is a minority view. Calling any architecture “the path to superintelligence” remains speculative on all sides. This is worth watching closely; it is not worth treating as settled.

Three Implications

IMPLICATION 1 — Continual Learning Is the Missing Primitive

The Skyfall benchmark made the deficiency measurable; Oak is the first pure-play research lab organized entirely around closing it. If continual learning at inference time is the capability that unlocks the next generation of useful agents, then the lab that owns that primitive owns a structural position no amount of fine-tuning or retrieval-augmented generation can replicate. Watch for enterprise AI buyers — especially in dynamic, fast-changing environments like financial trading, robotics, and autonomous logistics — to treat this research thread as a strategic dependency within 24 months.

IMPLICATION 2 — The Capex Assumption Has a Named Critic With Standing

Until now, the loudest skeptics of the compute supercycle were economists and short-sellers. Sutton is something different: a 2024 Turing Award winner with a specific, testable alternative hypothesis and a lab organized to prove it. If Oak publishes results — even early, partial results — showing that a continual-learning RL agent beats a same-generation LLM on adaptive tasks with less compute, that is a data point that moves infrastructure investment narratives. It does not need to be AGI to matter; it needs to be real and reproducible.

IMPLICATION 3 — The Neolab Era Is a Structural Bifurcation, Not a Blip

SSI, Thinking Machines, Oak, and the LeCun world-model cluster are not failed OpenAI applications. They are a deliberate architectural fork. Capital is now bifurcating: frontier-scale investments for the LLM incumbents, and smaller but serious bets on the labs building the architecture that the incumbents cannot pivot to without dismantling their core product. The War of Agents Architectures and The Agentic AI Stack are the right frames for tracking how this fork resolves over the next three to five years.

Business Engineer Framework

The Map of AI — Where Oak Lab Sits in the Stack

The Map of AI tracks 200+ companies across 9 layers of the AI stack — from infrastructure and compute through model development, agent architectures, and application deployment. Oak Lab represents a rare event: a new entrant attempting to rewrite the rules at the model and agent layers simultaneously, with a continual-learning primitive that, if validated, would force a reclassification of nearly every

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Sources: thelogic.co · oaklab.ai · x.com · en.wikipedia.org

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA