Based on the Vercel AI Gateway Production Index and a record single-day reading flagged on August 22, 2026.
A live record reading on the Vercel AI Gateway puts open-weight models at roughly 62% of token volume on August 22 — but the same data shows them capturing under 9% of the spend. The barbell is real, and it tells you exactly where commoditization stops today.
What Happened
The Vercel AI Gateway Production Index has been tracking one of the cleanest windows into real routing behavior available: which models developers and startups actually send tokens to, and what they pay for them. The published monthly figures show a clean directional trend — open-weight token volume on the gateway ran from 11% in April to 29% in June to 36% in July. On August 22, a live reading flagged a single-day record of roughly 62%, up from approximately 28% two months earlier. That figure is not in any published monthly index, it is one gateway on a Saturday when traffic skews toward experimentation over production, and single-day readings are noisy. But it is the continuation of a real, published trend — not a blip.
Read the number correctly, though, because the headline figure and the actual business story point in different directions. Open-weight models may be winning the majority of token volume on this gateway, but they are capturing under 9% of the spend. In the published July data, open weights ran 36% of tokens on 8.6% of dollars — itself a doubling from under 4% of spend in June, with more than 90% of that spend growth attributable to Moonshot’s Kimi K3 and Z.ai’s GLM. Anthropic, meanwhile, held 65% of all gateway spend on 30% of the volume in July, at 4.4 times the average token price on the gateway. In June that was 61% of spend on 32% of volume. Anthropic’s spend share is growing even as its volume share shrinks.
The engine behind the volume surge is specifically the Chinese open-weight wave. DeepSeek is now the second-largest source of tokens on the gateway at roughly a quarter of the total — more than twice Google’s approximately 11%. Kimi K3, released July 16, burns approximately 12 times the tokens per request that its predecessor K2.5 did: it is an agentic model built for long-horizon work, and token-heavy requests are structurally what long-horizon agentic tasks produce. The average price paid per token across the entire gateway fell 13.6% in July alone, a direct consequence of this volume mix shift.
The key insight: The honest read on the ~62% figure is not that open weights have won — it is that buyers are now routing the majority of their token volume to open weights while deliberately keeping the expensive, high-stakes work on the closed frontier. Routing discipline, visible in aggregate. The money — two-thirds of gateway spend at a 4.4× price premium — is sitting exactly where buyers are least willing to get things wrong.

The Structural Read
For two years, the case for open-weight models was supply-side: the models existed, they were free to run, and their quality was closing the gap with closed frontier labs. What the Vercel gateway data makes visible for the first time is the demand-side confirmation: buyers are actually routing the majority of their token volume to open weights. That is a different and more durable signal. Supply creates optionality; demand creates gravity.
But the barbell that emerges from the data is the more important structural fact. The market is not choosing open weights over closed frontier models — it is sorting work between them with increasing discipline. Cheap, high-volume, lower-stakes tasks — batch processing, content generation, experimentation, agentic subtasks — flow to open weights, where the marginal token cost approaches zero. High-value, high-risk work — customer-facing inference, complex reasoning, anything where a failure carries real cost — stays with Anthropic and the other frontier labs, which is why one provider can hold 65% of spend on 30% of volume at 4.4× the average price. This is what commoditization as a demand-side fact looks like: not a takeover, but a barbell, and the price data confirms it — average token price on the gateway down 13.6% in a single month.
The engine of that volume surge is specifically the Chinese open-weight cohort. DeepSeek at roughly a quarter of all gateway tokens — more than twice Google’s share — plus Moonshot’s Kimi K3 and Z.ai’s GLM, both agentic models designed for long-horizon tasks that structurally burn far more tokens per request than single-turn predecessors. This is the demand-side mirror of the supply-side story: Alibaba open-weighting Qwen is an explicit feed-the-ecosystem strategy, and what shows up on a routing layer like Vercel’s is where that strategy lands in actual usage. The Map of AI framing makes this legible: the model layer is the most contested, and open weights are now a credible demand-side alternative for the majority of token work — just not for the majority of token value.
BE Framework — The Harness Unlock
The constraint is no longer the model. It is the tooling built around a default frontier model.
Switching from a closed frontier model to an open one is only cheap if the CLIs, IDEs, SDKs, and agent harnesses around the model are model-agnostic. Most of the developer toolchain was built with a default frontier model baked in. As those tools go model-agnostic — as the harness layer adapts — the switching cost falls and the open-weight flywheel compounds into the spend column, not just the volume column. Until that happens, the open-weight story is a distribution and loss-leader play. The AI value chain still routes the money through whoever owns the harness.
One forward-looking caveat deserves explicit treatment. The thesis that harness adaptation will let the volume shift compound into a spend shift is analysis — a reasonable extrapolation from the trend, not a measured fact. The tooling could adapt slowly. The frontier labs could cut prices fast enough to defend volume; one of them just did. And the Vercel gateway, reflecting a developer- and startup-heavy population, underrepresents enterprises with large committed contracts to a single frontier provider — exactly the traffic that would pull the spend number toward closed models. The ~62% is a live Saturday reading on one gateway. The published monthly index had open-weight volume around 29–36% through July. Both facts are true, and neither one should be discarded to make the other read cleaner. See also: Beyond NVIDIA’s Moat for the full stack framing. The routing-layer dynamic is also covered in the Ramp Router and routing-war analysis.









