Vercel AI Gateway Leaderboard: DeepSeek Dominates Token Volume, Claude Dominates Spend

On Vercel’s AI Gateway leaderboard, one model runs 59.3% of the tokens and collects 5.1% of the spend. Another runs 1.7% of the tokens and collects 13.7% of the spend. Both numbers are on the same page — and together they expose a structural split that a single “who’s winning” question cannot resolve.

VERCEL AI GATEWAY — ONE GATEWAY’S TRAFFIC · 21 JUN – 18 SEP 2026

59.3%

DeepSeek V4.1 Flash — token volume share

13.7%

Claude Opus 4.8 — spend share

5.1%

DeepSeek V4.1 Flash — spend share

1.7%

Claude Opus 4.8 — token volume share

78.4%

Open-weight models — share of token volume

15.6%

Jev — Reach, the board’s widest-adoption metric

This is Vercel AI Gateway traffic — a real measurement of a selected developer population, not market, industry or global share. Not investment advice.

What Happened

Vercel’s AI Gateway publishes daily-updated leaderboards of the traffic passing through it, and the board covering 21 June to 18 September 2026 shows four separate ranked lists on the same page: token volume, spend, preference, and reach. This is Vercel AI Gateway traffic — a real, anonymised observation of a selected population of developers who have chosen to route requests through that gateway — and not a measurement of the broader market or industry.

On the token-volume board, DeepSeek V4.1 Flash leads at 59.3%, followed by GLM 5.3 Flash at 7.5%, GPT 5.6 Luna at 4.0%, DeepSeek V4 Flash 0731 at 2.7%, Kimi K3 at 2.5%, Muse Spark 1.3 Contributor at 2.0%, Claude Opus 4.8 at 1.7%, and Claude Sonnet 5 at 1.7%, with 16.5% in other models. Models that publish their weights for download account for 78.4% of token volume on this board against 21.6% for everything else — the page’s own definition and its own figure.

On the spend board, the ranking inverts almost completely: Claude Opus 4.8 leads at 13.7%, then GPT 5.6 Sol at 9.5%, Kimi K3 at 9.3%, Claude Opus 5 at 9.0%, GPT-6 Astra at 7.0%, Claude Sonnet 5 at 5.5%, DeepSeek V4.1 Flash at 5.1%, Claude Opus 4.6 at 4.4%, and Claude Sonnet 4.6 at 4.3%, with 32.2% in other models. GPT 5.6 Luna and GPT 5.6 Sol are separate entries on the board and are treated as such throughout this piece.

The key insight: Volume share and spend share are different quantities, and on this board they have come apart almost completely. DeepSeek V4.1 Flash does roughly thirty-five times Claude Opus 4.8’s share of the token work and receives roughly a third of its share of the spend — and both figures are published by the same source on the same screen. The question “who is winning” is underdetermined until you specify which quantity you mean.

A unit of work and a unit of value are not the same unit. On this board they rank the field in opposite orders
A unit of work and a unit of value are not the same unit. On this board they rank the field in opposite orders, and no figure on it is wrong.

The Token-Volume Board: What Ran

DeepSeek V4.1 Flash 59.3%
GLM 5.3 Flash 7.5%
GPT 5.6 Luna 4.0%
Kimi K3 2.5%
Claude Opus 4.8 1.7%

Token volume share, Vercel AI Gateway, 21 Jun – 18 Sep 2026. Remaining share distributed across other models.

The Spend Board: What Was Paid For

Claude Opus 4.8 13.7%
GPT 5.6 Sol 9.5%
Kimi K3 9.3%
Claude Opus 5 9.0%
GPT-6 Astra 7.0%
DeepSeek V4.1 Flash 5.1%

Spend share, Vercel AI Gateway, 21 Jun – 18 Sep 2026. The board’s “Other” bucket is 32.2%; further named entries not shown here are Claude Sonnet 5 at 5.5%, Claude Opus 4.6 at 4.4% and Claude Sonnet 4.6 at 4.3%.

The Structural Read

The easy reading of this board is that open-weight models dominate volume because they are the cheap ones, and that expensive proprietary models capture the spend. That reading is visible in the data in a broad sense — models carrying tier labels like “Flash” do populate the top of the volume board, and open-weight models account for the majority of token volume on this gateway. But the board itself refutes the simple version of the story.

Kimi K3 publishes its weights. It ranks third in spend at 9.3% on only 2.5% of token volume. An open-weight model is therefore sitting near the top of the money board — which the licence-based explanation cannot accommodate. The more careful reading is that the two boards separate high-throughput work from high-value work, and that this separation corresponds to model tier rather than to whether weights are published. Open weights and low cost overlap substantially on this board without being the same category, and the distinction matters: one is a property of a licence, the other is a property of a product line.

Structural Observation

“A leaderboard can rank a field in one order by volume and close to the opposite order by spend without any figure on it being wrong. A reader who takes either board alone as the answer has not been misled by the data so much as by the question.”

The preference and reach boards add a third and fourth answer to the same question. On preference — defined as the share of gateway teams using a model as their primary model — the leader is Jev at 13.3%, ahead of GPT 5.6 Luna and Claude Sonnet 5 at 6.9% each, Claude Haiku 4.5 at 6.7%, and Claude Sonnet 4.6 at 4.2%. On reach, Jev leads again at 15.6%, ahead of Claude Haiku 4.5 at 14.6%, GPT 5.6 Luna at 13.0%, and Claude Sonnet 5 at 12.2%.

Jev, as reported, has no open-weight release and no announced self-hosting path, and is available through a hosted interface behind an early-access waitlist. The structural observation is that volume measures what ran, spend measures what was paid for, and preference measures what a team standardised on — three different decisions, made by different people, at different moments, under different constraints. Nothing requires them to agree.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

This is not investment advice. Every figure above describes traffic through a single gateway — Vercel’s AI Gateway — over 21 June to 18 September 2026, and is not market, industry or global share. It is an exact measurement of a selected population of developers and is silent about every other population; nothing above extrapolates to the industry or speculates about traffic elsewhere. GPT 5.6 Luna and GPT 5.6 Sol are separate entries on the board and are not treated above as the same model. Kimi K3 publishes its weights and ranks third on the spend board, which is why nothing above argues that open-weight models are simply the cheap ones. That Jev has no open-weight release and no announced self-hosting path is as reported; nothing above describes its architecture, pricing, benchmarks or quality. No price per token, absolute token, request or dollar figure, laboratory revenue, margin or quality ranking appears above, and nothing above says why any team chose any model or which of these metrics is the correct one. Nothing is predicted.

Sources: vercel.com · techcrunch.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA