Cognition Runs Vera Rubin NVL72 on CoreWeave

Nvidia’s newest rack is live on CoreWeave with one named customer, four multipliers, and four different denominators — and only one of them tells you what you need to know.

The 4.8x and 3.8x figures below are Cognition’s own, measured by Cognition on CoreWeave against CoreWeave’s GB200 NVL72 baselines and published by CoreWeave. They are per GPU, at matched interactivity, on Cognition’s SWE-2 model. Three primaries were read for this piece: CoreWeave’s engineering posts of 30 September and 21 July 2026, and Nvidia’s own blog post of 30 September. All three are vendor material, and no independent party has verified any figure. Nothing here is investment advice.

What Happened

CoreWeave announced on Wednesday, September 30, 2026, that Nvidia’s Vera Rubin NVL72 is in limited availability on its cloud. Cognition — maker of the Devin AI coding agent — is the first customer running production workloads on it. CoreWeave’s engineering blog, authored by Harsh Singh Banwait, is one primary source. Nvidia published its own account the same day, written by Stuart Pitts, and it carries detail CoreWeave’s post leaves out.

The headline figure from the announcement is 4.8 times the total token throughput per GPU of an Nvidia GB200 NVL72, at matched interactivity, on Cognition’s own SWE-2 inference workload. That qualifier matters. The post’s opening paragraph drops it, stating only that Cognition “is seeing a 4.8x increase in total token throughput.” The precise version appears thousands of words further down.

CoreWeave completed first bring-up and validation of the Vera Rubin NVL72 platform in early June 2026. It published the first measured silicon numbers on July 21, 2026, using DeepSeek R1 with expert parallelism, NVFP4, multi-token prediction, and disaggregated prefill and decode enabled. Wednesday’s post is the first production customer announcement.

The two per-GPU figures and the per-megawatt figure are quoted at matched interactivity against GB200 NVL72. T
The two per-GPU figures and the per-megawatt figure are quoted at matched interactivity against GB200 NVL72. The 45x token-cost claim appears in the announcement with no methodology attached, which is why it sits at the top of this chart and should carry the least weight.

The Structural Read

The Product Overhang Doctrine says capability builds invisibly until it surfaces all at once. That is the honest description of what CoreWeave is doing here. The platform exists at limited scale — hundreds of Rubin GPUs, across multiple unnamed regions, in one 72-GPU-per-rack configuration — but the numbers it is posting are designed to define the category before general availability arrives.

The phrase doing the real work in Wednesday’s post is “at matched interactivity.” In agentic inference, every step waits on the one before it. Per-step responsiveness is the binding constraint, not raw throughput. A throughput figure quoted without an interactivity target is not a usable number for an agentic workload.

That is why Cognition is the right first customer for this announcement. Silas Alberti, SVP Research and Founding Team at Cognition, explains the compounding directly.

Silas Alberti — SVP Research, Cognition

“For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done.”

The benchmark provenance is also worth stating plainly. Cognition’s engineering team conducted the benchmarks on CoreWeave’s infrastructure, comparing Vera Rubin NVL72 against GB200 NVL72 baselines. The customer ran the test, on the vendor’s cloud, on its own SWE-2 model, and the vendor published the result. That is not third-party validation. What the workload actually was appears only in Nvidia’s version of the announcement, not CoreWeave’s: Cognition sampled a subset of tasks from FrontierCode and deployed agents to solve them. SWE-2 is the model doing the solving. FrontierCode is where the problems came from.

All of this is disclosed in the post, which is to CoreWeave’s credit — it is more than most infrastructure announcements offer. Disclosure does not convert a first-party number into an independent one.

There is also a discrepancy inside CoreWeave’s own published material. Wednesday’s post gives the platform 216 TB/s of NVLink bandwidth. The July 21 post, by the same author, describes a 260 TB/s all-to-all NVLink 6 fabric. Those two figures do not reconcile from the two posts alone. Neither explains the other. Both are reported here as published.

What It Changes

Three things follow from the way this announcement is written, and none of them depends on taking the numbers on trust.

The first is that interactivity has become the axis. A raw throughput figure with no latency anchor tells an agentic buyer nothing, because the workload is a chain of steps and each one waits on the last. CoreWeave quotes every measured figure at matched interactivity, which is a more careful standard than most accelerator announcements meet.

The second is that four multipliers with four denominators is a reader problem. The 4.8x is per GPU at matched interactivity on SWE-2. The 10x is per megawatt on DeepSeek R1, from a different post two months earlier. The 45x carries no methodology at all. Nothing in the material invites a reader to combine them, and nothing stops one either.

The third is that limited availability is not a capacity story. Hundreds of Rubin GPUs is a handful of 72-GPU racks. The architecture is real. CoreWeave describes 72 Rubin GPUs and 36 Vera CPUs presented as one logical unit over NVLink 6, with 1,400 TB/s of HBM4 memory bandwidth. The operational questions that would make it matter at scale — price, regions, a general-availability date — are all unanswered in the post. This is a performance claim, not a supply event.

Business Engineer Framework

Product Overhang Doctrine

The Product Overhang Doctrine maps how capability accumulates below the surface before it becomes a market event. Vera Rubin’s limited availability is exactly that phase: real performance, constrained supply, no public pricing. The Map of AI traces where infrastructure announcements like this one sit across all nine layers of the AI stack — and which layer captures the value when the overhang finally releases.

Explore the Map of AI →

The Bottom Line

Vera Rubin NVL72 is delivering 4.8 times the inference token throughput per GPU of a GB200 NVL72, at matched interactivity, on Cognition’s SWE-2 workload — and 3.8 times on reinforcement-learning runs under the same condition. Those are first-party numbers from a customer-run benchmark on vendor infrastructure, disclosed fully and therefore worth taking seriously but not treating as independent. The fleet is hundreds of GPUs across a handful of racks.

The 45x token-cost figure has no published methodology. Two NVLink bandwidth figures from the same author — 216 TB/s and 260 TB/s — sit unreconciled across two posts. The architecture is real. The scale is not yet.

Sources: CoreWeave Engineering Blog — Cognition First Customer Announcement (30 Sep 2026); CoreWeave Engineering Blog — 10x Tokens Per Megawatt (21 Jul 2026). Both CoreWeave posts authored by Harsh Singh Banwait. Nvidia’s own post of 30 September 2026, by Stuart Pitts, was also read: From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Both primary sources for this piece are CoreWeave’s own engineering blog — the post of 30 September 2026 announcing limited availability, and the post of 21 July 2026 reporting first measured silicon performance. Both are by the same author. Nvidia published its own post on the same deployment the same day, written by Stuart Pitts, and it was read for this piece. All three are vendor material and no independent party has verified any figure here.

The two accounts are not identical: the task source for the benchmark, FrontierCode, appears only in Nvidia’s post, and the two companies quote Silas Alberti saying different things on the same day. The 4.8x and 3.8x results are Cognition’s. Cognition ran the benchmarks on CoreWeave, against CoreWeave’s GB200 NVL72 baselines, using Cognition’s own SWE-2 model, and CoreWeave published the outcome.

CoreWeave discloses that arrangement in the post, which is more than most infrastructure announcements do. It remains a first-party result rather than third-party validation. The four multipliers in that material are measured against four different things: inference tokens per GPU, reinforcement-learning output tokens per GPU, tokens per second per megawatt on a different model, and token cost. The first three are quoted at matched interactivity.

The 45x token-cost figure appears with no methodology and no link, and is reported above on that basis. The 216 TB/s and 260 TB/s NVLink figures are both quoted as published. Nothing above decides between them, because the two posts do not reconcile and neither explains the other. Not established and therefore absent: price per GPU-hour, rack count, which regions, any general-availability date, and the absolute tokens-per-second numbers behind either ratio. Nothing above predicts anything, and nothing here is investment advice.

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA