Context via Fortune and The Next Web.
Moonshot AI’s Kimi K3 revived the fear that capable, cheaper AI would gut chip demand and strand $700 billion in hyperscaler capex — but the history of technology suggests the causality runs the other way.
What Happened
When Moonshot AI released Kimi K3 in mid-July 2026, the narrative that formed around it was familiar. Here was a large, open-weight model delivering near-frontier capability at roughly $15 per million output tokens — against approximately $25 to $30 per million for the leading US systems, Claude Opus and GPT-5.6. That is a meaningful discount: about 1.5 to 2 times cheaper, not a wholesale price collapse, but real enough to reignite a question that first surfaced with force about eighteen months earlier, during what markets labeled the DeepSeek moment in January 2025. The question: if capable AI keeps getting cheaper — and open-weight models can be downloaded for free — does the roughly $700 billion that hyperscalers are committing to AI infrastructure ever pay back?
Chip stocks sold off partly on that logic. The implied mechanism is straightforward: lower price per token → lower revenue per unit of compute → less reason to build more compute. It is a tidy story, and it has the causality exactly backwards. Kimi K3 is a genuine technical achievement — its performance on coding and reasoning benchmarks puts it in genuine contention with closed US frontier models at a lower price point. But “capable AI is getting cheaper” and “demand for compute is falling” are not the same statement. Conflating them is how the delusion works.
It is worth being precise about what Kimi K3 actually represents in the pricing landscape. At roughly $15 per million output tokens, it is cheaper than the US frontier — but it is not the cheapest model on the market. This is a story about the broad, steady decline in AI inference costs as open-weight capability catches up with closed-lab performance, not a single-event price collapse. That distinction matters for how you reason about the demand implications.
The key insight: Every prior technology that got materially cheaper — coal, bandwidth, cloud compute — saw total consumption rise rather than fall. The price decline expanded the population of viable use cases faster than it compressed per-unit revenue. There is strong structural reason to expect AI inference to follow the same path, and early platform data on open-weight model usage is already pointing in that direction — though the pattern is a tendency, not a guarantee.
The Structural Read
The economic principle at work here is the Jevons paradox, named for the nineteenth-century economist William Stanley Jevons, who observed that improvements in coal-burning efficiency led to a rise in total coal consumption rather than a fall. The mechanism is not mysterious: when a resource becomes cheaper to use per unit, applications that were previously uneconomical become viable, entirely new categories of use emerge, and adoption spreads to actors who could not previously afford to participate. The net effect on total consumption is almost always positive, often dramatically so. It happened with coal, with telecommunications bandwidth in the 1990s and 2000s, and with cloud computing as AWS drove down the cost of provisioning infrastructure.
Applied to AI inference, the logic runs as follows. At $25 to $30 per million output tokens, many AI use cases — high-volume document processing, always-on customer interaction, real-time coding assistance at scale — carry cost structures that make them difficult to justify at production volumes. At $15 per million, or lower, some of those cases become viable. More companies can afford to experiment. Developers who were previously rate-limiting their API calls for cost reasons stop doing so. Workloads that did not exist at $30 get built at $15. The OpenRouter platform data on Chinese open-weight model adoption is already pointing in this direction: the rise of models like DeepSeek and Kimi in developer usage appears to be largely additive — new workloads coming online — rather than a clean substitution away from paid frontier APIs. That evidence is qualitative and directional rather than definitive, but it is consistent with the Jevons mechanism.
The honest caveat belongs in the same sentence as the claim: Jevons is a tendency, not a law. If model capability reaches a genuine plateau — where further price declines produce no meaningful expansion in viable use cases — then cheaper could compress spending in the near term rather than expand it. If aggregate demand saturates before new use cases scale, the arithmetic changes. And even when total compute volume grows, that does not guarantee that incumbent infrastructure providers capture the value; a cheaper commodity inference layer can shift margin to application builders while volume accrues at the hardware level with thinner economics. The payback risk on $700 billion of hyperscaler commitments is real, and the timing is genuinely uncertain. The Jevons argument addresses the direction of compute demand, not the distribution of who profits from it.
The Kimi Delusion
The bearish take has the causality inverted
Cheaper capable models do two things simultaneously: they expand the addressable demand for inference compute, and they force closed frontier labs to differentiate forward — into harder reasoning, longer context, and agentic capability — which requires more training compute and more inference capacity, not less. An open-weight price shock does not remove the reason to build data centers. It raises the bar that closed labs must clear, and clearing that bar costs capex. The market’s own capacity allocators made this judgment explicit: TSMC and ASML raised guidance and expanded capacity plans the same week the cheap-AI fear peaked.
This is the second half of what the Business Engineer essay on the Kimi Delusion develops in full. The structural pressure that open-weight models place on closed labs is not a demand-destruction event — it is a competitive forcing function. OpenAI, Anthropic, and Google cannot respond to a capable $15 model by standing still. They respond by investing in the next capability tier: better reasoning chains, longer context windows, deeper tool use, more reliable agentic workflows. Each of those responses is compute-intensive, both in training and in inference. The treadmill accelerates rather than slows. This dynamic sits within the broader capital-intensity logic explored in The Subsidized AGI Economy — a framework for understanding why the frontier AI race continues to attract capital even as per-token economics decline.
Three Implications
COMPUTE DEMAND — DIRECTIONALLY HIGHER
The Jevons mechanism suggests that steadily declining inference costs will expand the population of economically viable AI workloads faster than they compress per-unit revenue at the infrastructure layer. Open-weight adoption on routing platforms is already behaving additively. The near-term caveat is real — if capability plateaus or demand saturates, the direction could reverse — but the structural prior favors rising total compute consumption, not falling.
CLOSED-LAB STRATEGY — FORCED UPMARKET
Every wave of capable open-weight releases compresses the defensible market for mid-tier closed models and pushes frontier labs to invest harder in the next capability tier. That is not a threat to their capex rationale — it is the rationale. The companies that can credibly deliver on harder reasoning, agentic reliability, and enterprise integration maintain pricing power; those that cannot will face genuine margin pressure even if total compute grows.
VALUE CAPTURE — THE OPEN QUESTION
Higher total compute volume does not automatically mean higher margins for any particular layer in the stack. A commoditizing inference layer shifts bargaining power toward application builders and away from model providers and, potentially, toward hardware vendors who benefit from volume regardless of who runs the models. Investors and strategists need to track not just whether compute demand grows, but where in the stack the economics accrete — that question remains genuinely open.
The Bottom Line
Kimi K3 is a real model making a real dent in the price of near-frontier AI inference, and the market’s reflexive selloff on cheaper AI is a real mistake — not because the payback and timing risks on $700 billion of hyperscaler capex are imaginary, but because the mechanism behind those fears has the causality inverted. Cheaper capable AI expands the universe of economically viable workloads, forces closed labs to invest harder in the next capability tier, and pushes total compute demand upward; the open-weight evidence on routing platforms is already directionally consistent with that read. The honest position holds the Jevons hedge — the tendency is not a law, value-capture can shift even as volume grows, and near-term compression is possible if capability stalls — but it does not mistake a price decline for a demand collapse. Those are different claims, and conflating them is exactly how the delusion spreads.
Sources:
Fortune — Kimi K3 launch and market reaction ·
The Next Web — Kimi K3 and the tech selloff ·
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.









