Vera Rubin’s 67x Gain Is 1.4x–3x at 60–100 TPS

These are SemiAnalysis’s own benchmark results on early pre-release software, produced with NVIDIA’s help on bring-up and verification, as the post says. This publication read only the free portion of the post and ran nothing.

SemiAnalysis’s Vera Rubin post gives two multiples that look like they disagree: about 67x the throughput per TCO of GB300 at 170 TPS, and 1.4x to 3x at 60 to 100 TPS, where it says most providers would serve. The results are SemiAnalysis’s own, on early pre-release software, produced with NVIDIA’s help. This publication read the free portion, ran nothing, and did not read the paywalled remainder.

The Two Numbers

The post, published 14 September 2026, reports what SemiAnalysis calls the first verified agentic inference results for Rubin, measured on its AgentX benchmark. It says even on early pre-release software the results show why extreme co-design was necessary.

In its charts the Y-axis is total tokens per $1 of TCO. The figures here use the owning assumption: the modeled cost to own and operate the hardware at large hyperscaler volume, including capex, colocation, power and cost of capital.

At 170 TPS, SemiAnalysis says Vera Rubin NVL72 delivers about 67x the total throughput per TCO of GB300 Dynamo TRTLLM. At the part of the frontier where it says most providers would actually serve the model, 60 to 100 TPS, the gain is between 1.4x and 3x against the latest GB300 TRTLLM configuration.

The same Rubin result is 5.56x or 62.9x at 170 TPS depending on which GB300 engine is the comparison. SemiAnal
The same Rubin result is 5.56x or 62.9x at 170 TPS depending on which GB300 engine is the comparison. SemiAnalysis’s own line: the engine label is essential when quoting the high-interactivity gain.

Same Benchmark, Different Points on the Curve

The two numbers answer different questions. The 67x sits at a high interactivity target, where the GB300 TRTLLM curve is near its fastest measured endpoint. The 1.4x to 3x covers the speeds at which, in the post’s words, providers would actually serve this model.

The post also says Vera Rubin reaches about 61% higher maximum P90 interactivity than GB300 Dynamo TRTLLM, 276.24 versus 171.53 P90 TPS. It adds that with the open-source SGLang stack, GB300 can reach similar interactivity to Vera Rubin.

Per Megawatt, the Engine Matters

For the DeepSeek V4 Pro agentic workload, SemiAnalysis says the size of Rubin’s advantage depends on the interactivity target and the serving engine. At 100 TPS it gives Rubin about 59.4 million total tokens per second per megawatt, against 28.5 million for GB300 Dynamo SGLang and 21.1 million for GB300 Dynamo TRTLLM. That is 2.09x over the stronger GB300 engine.

At 150 TPS the post gives Rubin nearly 37 million, about 7.2x GB300 SGLang. At exactly 170 TPS it gives 62.9x against GB300 TRTLLM and 5.56x against GB300 SGLang. The post’s own conclusion: the engine label is therefore essential when quoting the high-interactivity gain. The values are interpolated between measured points.

What the Podcast Clip Says

On SemiAnalysis’s Episode 034 podcast, Cam Quilici says that in the range of interactivity where providers would actually be serving DeepSeek V4 Pro, Rubin is “just like 3x better.” The clip names no metric or serving engine, so this publication does not match it to any row of the post. It read the clip’s transcript, not the benchmark data behind it.

What NVIDIA’s Post Says

NVIDIA’s blog of 1 October 2026 says SemiAnalysis AgentX data shows Vera Rubin NVL72 delivers over 30x higher throughput per megawatt than GB300 NVL72, and up to 45x lower cost per million tokens on DeepSeek V4 Pro. That sentence states no interactivity target or serving engine. This piece does not say which row of SemiAnalysis’s table it matches.

The Modelled Economics

SemiAnalysis also models revenue per gigawatt. At 75 TPS, 60% utilization and no model-license fee, it gives Vera Rubin $159.5 billion in annual revenue and $149.9 billion in modeled profit per all-in utility GW. For the strongest GB300 configuration, Dynamo SGLang, it gives $114.9 billion and $105.3 billion. That is about 39% more revenue and 42% more modeled profit.

What Is Not Established

Not established, and therefore absent from this analysis: any independent replication; results on mature software, since the post says performance should improve; the paywalled lifecycle revenue analysis; the cost per million tokens behind NVIDIA’s 45x; which row of the table NVIDIA’s 30x refers to; and any comment from NVIDIA beyond its blog sentence.

None of the above is investment advice. It reports what a benchmark publisher says about its own results.

Business Engineer Framework

Founders, Distributors, Enablers

The Business Engineer framework maps how value is created across an AI ecosystem into Founders, Distributors and Enablers. The Map of AI shows where each player sits.

Explore the Map of AI →

Every figure above comes from SemiAnalysis’s newsletter post of 14 September 2026, read from the free portion in the newsletter’s feed; the rest of the post is behind a paywall and was not read. They are SemiAnalysis’s own benchmark results (InferenceX and AgentX) on early pre-release software, produced with NVIDIA’s help on bring-up and verification, as the post says. This publication ran nothing and tested nothing, and found no independent replication in the sources it read.

The NVIDIA figures are one sentence from NVIDIA’s blog post of 1 October 2026. That sentence states no interactivity target or serving engine, and this piece does not say which row of SemiAnalysis’s table it matches. The revenue and profit per gigawatt figures are SemiAnalysis’s modelled numbers under its stated assumptions. Nothing above predicts anything and nothing here is investment advice.

Sources: newsletter.semianalysis.com · blogs.nvidia.com · newsletter.semianalysis.com — ‘Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar’, 14 September 2026 (free portion) · NVIDIA blog, 1 October 2026 — ‘Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment’ · SemiAnalysis podcast Ep. 034 clip (Cam Quilici)

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA