Based on Skyfall AI’s benchmark, “LLMs Are Not Continual Learners” (2026).
Skyfall AI’s warehouse benchmark finds that frontier models hold stable reward curves not because they adapt — but because they’re running frozen heuristics. Flat is not adaptive. And that distinction breaks the core assumption of the agent economy.


What Happened
Skyfall AI — a supply-chain AI company — published a benchmark bluntly titled LLMs Are Not Continual Learners, placing GPT-5.5 and Gemini 3.1 Pro inside a stateful warehouse-allocation environment: routing resources across queues while the system absorbs WMS capacity drops and EDI failure spikes. The models received reward feedback as conditions shifted. The question was simple — do they actually learn from it?
The headline result looks reassuring. GPT-5.5 posts Task-1 rewards of ~0.920, 0.920, 0.914 straight through the configuration changes. Gemini 3.1 Pro holds ~0.899, 0.853, 0.852. Stable. Consistent. Apparently competent. But Skyfall’s argument is that this stability is the warning, not the all-clear. The environment changes. The reward stays flat. And there is no evidence of any policy update. The model isn’t adapting — it’s coasting on whatever heuristic it arrived with.
The benchmark isolates two distinct frozen strategies. Gemini 3.1 Pro plays a concentration heuristic: fewer queues per step (mean ~1.34), high max-capacity share (~0.91), allocation score ~0.864 out-of-distribution and ~0.948 in-distribution. GPT-5.5 plays the mirror: a diversification heuristic across more queues (~2.11–2.33 per step), lower max-capacity share (~0.74–0.85), allocation ~0.918 out-of-distribution and ~0.885 in-distribution. Both are strategies a competent human logistics manager might plausibly choose. Neither was learned from the reward signal. Neither closes the gap to the reward-optimal upper bound.
The key insight: A flat reward curve in a changing environment is not evidence of competence — it’s evidence of a frozen prior. The model that never gets worse also never gets better. On day 300 of a live deployment, it will be exactly as (mis)calibrated as it was on day 1.
The Structural Read
The agentic-AI thesis rests on a load-bearing assumption: deploy an agent and it will get better on the job. Skyfall’s benchmark — and it must be read with that caveat front and center, since Skyfall is a supply-chain AI vendor with an obvious interest in framing where general LLMs fall short — attacks that assumption with unusual precision. The scope is narrow: two models, one warehouse domain, and “continual learning” defined specifically as in-context policy adaptation from reward feedback, not weight updates. But the logical structure of the argument is clean enough to generalize.
Three structural cracks appear. First, stable performance is a mirage. A flat reward line gets read as “it works” on every ops dashboard. But if it’s a frozen pretrained heuristic, the agent will be exactly as good — or bad — on day 300 as on day 1, and silently mediocre whenever conditions drift outside the distribution it was pretrained on. “It hasn’t gotten worse” is not “it’s learning.” Second, learning is a missing layer, not a model property. If frontier models don’t close the gap to reward-optimal on their own — and the benchmark shows they don’t — then adaptation has to come from outside the model entirely.
This is exactly the insight that connects to Satya Nadella’s formulation about enterprise AI moats: the durable advantage isn’t the model, it’s the learning loop around it — your evals, your traces, your adapted weights. Sierra’s argument that context beats raw capability reads as the same point from the product side. And the RL-data labor infrastructure that Mercor priced at a $20B valuation is precisely the scaffolding that has to sit between raw model and self-improving agent. Claims about self-improving agents like Perplexity Brain should be measured against exactly this bar: is the reward gap actually closing, or is the reward curve just flat?
Skyfall AI Benchmark
“Flat is not adaptive. The reward stays steady because the model is applying the same fixed heuristic across changing conditions — not because it’s updating its policy.”
The third crack is the most operationally damaging: opaque failure blocks enterprise autonomy. When GPT-5.5’s Task-2 reward collapses to zero — repeatedly, unpredictably — context overflow, a pre-training coverage gap, and reward misalignment all look identical from the outside. An RL failure maps to a diagnosable mechanism. An LLM failure is a behavior pattern with no address. If you can’t attribute a dropped reward, you can’t fix it. And if you can’t fix it, a human has to stay in the loop to catch it — which caps how far “agentic” can go in high-stakes operations. Autonomy stalls exactly at the point where diagnosability runs out.
Three Implications
IMPLICATION 1 — “Stable” SLAs Are Not Learning Proof
Any vendor showing flat, stable reward curves as evidence of agent performance needs to answer a harder question: is that stability the result of genuine adaptation, or a frozen heuristic that happens to be in-distribution today? The benchmark gives enterprise buyers the right question to ask — and most current agent demos don’t have an answer.
IMPLICATION 2 — The Moat Is the Loop, Not the Model
If frontier models don’t self-improve from reward, then the companies that build the RL fine-tuning layer, the eval infrastructure, the memory and trace stack — those are the durable competitive positions. The model is a commodity input. The continual-learning scaffold around it is the asset. This is precisely what the Agentic AI Stack and the Four Intelligence Moats predict as the next layer of competitive advantage.
IMPLICATION 3 — Diagnosability Is a First-Class Product Requirement
Opaque failure isn’t a bug to patch — it’s an architectural constraint that determines how much autonomy an enterprise will actually grant. The next unlock in agentic AI isn’t a larger model; it’s observability and failure attribution tooling that lets operators know whether a collapsed reward was a context issue, a distribution shift, or a misaligned objective. Without that layer, “agentic” is a deployment-level aspiration that human oversight permanently caps.
The Bottom Line
Skyfall’s benchmark — self-interested scope and all — delivers a precise stress test to the agent economy’s core claim. GPT-5.5 and Gemini 3.1 Pro are not getting better on the job; they arrived with heuristics, they’re running heuristics, and
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.









