NVIDIA VP Jason Hardy’s on-record argument to TechCrunch is strategic positioning first — but the frame it proposes, and what it concedes, tells you more than the numbers do.
What Happened
In on-record comments to TechCrunch published August 29, 2026, NVIDIA’s VP of storage technology Jason Hardy made an argument that is best understood as strategic positioning before it is taken as technical fact. Hardy’s claim: at gigawatt-scale AI infrastructure, the hard problem is not raw compute but moving data — memory limits per server and flash bottlenecks become the real ceiling. Offloading orchestration work to the Vera CPU produced, in NVIDIA’s telling, “upwards of 3x improvement in these operations.” That figure is NVIDIA’s own claim for a specific orchestration-offload operation on its own hardware; it has not been independently benchmarked.
The vehicle for this argument is Vera Rubin, which is currently rolling out and which NVIDIA is pitching not as a faster GPU but as a full stack — GPU, CPU, inference accelerators, storage, and networking sold as one integrated system. The pitch reframes the unit of competition from the chip to the rack. For Hardy, that is precisely the point.
Notably, TechCrunch’s own reporting — independent of Hardy’s comments — draws a parallel that gives the argument structural weight: OpenAI’s Jalapeno inference chip, co-developed with Broadcom, has a stated design goal of minimizing data movement and communication delays. That detail was surfaced by the reporter, not volunteered by Hardy, and Hardy’s interview does not mention Jalapeno, NVLink Fusion, Google TPU, or Amazon Trainium. The convergence — incumbent and its largest customer both pointing at data movement as the binding constraint — is the part worth sitting with.
The key insight: Hardy’s argument is an NVIDIA VP’s on-record framing that structurally favors NVIDIA — treat it as such throughout. What it concedes is more revealing than what it claims: when the world’s dominant GPU vendor argues publicly that FLOPs are no longer the fight, it is signaling that the die is now contestable, and racing to make the die not the point.
The Structural Read
The classic move of a leader who senses the ground shifting under its strongest asset is this: when you can’t be certain of winning the die, move the game to the system. That is precisely what Hardy’s argument does, and it is worth taking seriously on its own structural logic — separately from whether NVIDIA’s numbers hold up to scrutiny.
For two years, the structural threat to NVIDIA has been the custom-silicon bifurcation: hyperscalers designing their own accelerators that win on perf-per-watt for their specific workloads and gradually erode NVIDIA’s merchant-GPU margin. A well-funded single customer — Amazon, Google, Microsoft, OpenAI — can now tape out a chip that beats NVIDIA on its own inference workload. That is the contestable die. (For the Marvell and custom-silicon picture, see our deep-dive on Marvell’s FY28 custom-silicon strategy.)
Hardy’s counter is elegant in the way good strategic reframing always is. Redefine the product from the chip to the rack, and declare the bottleneck to be orchestration and memory movement rather than FLOPs. The moment that framing is accepted, a rival ASIC stops being a replacement for NVIDIA and becomes a component that plugs into NVIDIA’s system — specifically, into the interconnect and data-movement layer that NVIDIA controls. The moat migrates from the die, where a well-funded single customer can now match or beat NVIDIA, to the rack-scale interconnect and system-integration layer, where matching NVIDIA means replicating an entire vertically-integrated stack rather than taping out one good chip.
Jason Hardy, NVIDIA VP of Storage Technology — via TechCrunch, Aug 29 2026
“Upwards of 3x improvement in these operations.”
Context: NVIDIA’s own claimed figure for orchestration-offload operations via the Vera CPU; not independently benchmarked. Hardy’s full comments via TechCrunch.
The corroboration that this is more than spin comes from TechCrunch’s reporting itself, which draws the parallel to Jalapeno independently of Hardy. When both the incumbent and its largest customer point at data movement as the binding constraint at scale, it suggests the framing reflects a real engineering reality — even if NVIDIA’s preferred resolution of that constraint happens to favor NVIDIA. Both sides agree the fight has moved to the system. The question is whose system.
The logical endpoint of NVIDIA’s full-stack pitch is what NVLink Fusion, NVIDIA’s separately-announced program, makes explicit: absorbing rival silicon as XPUs inside NVIDIA racks, so that hyperscaler ASICs become components in NVIDIA’s architecture rather than replacements for it. That is NVIDIA’s preferred outcome and a sales motion — not a done deal. Whether hyperscalers accept that frame, or build their own rack-scale systems around their own ASICs, is the unsettled question this argument is designed to influence. Hardy’s TechCrunch comments are NVIDIA’s opening bid in that negotiation, not the result of it. For a deeper map of where the moat now sits, see Beyond NVIDIA’s Moat at Business Engineer.
BE Framework Read
Move the Moat From the Die to the Rack
When the incumbent can’t guarantee winning at the chip level, the playbook is: redefine the product boundary upward. Make the chip a component of a system you control, so the competitor’s win at the chip level becomes your win at the system level. The moat migrates to wherever a single good chip is insufficient — in this case, rack-scale interconnect, orchestration, and system integration. Matching NVIDIA then means replicating an entire vertically integrated stack, not taping out one competitive die. This is the structural logic of the full-stack pitch; it survives the discount on NVIDIA’s specific numbers.
Three Implications
IMPLICATION #1 — For Hyperscalers Weighing Custom Silicon
The full-stack pitch is NVIDIA’s attempt to talk hyperscalers out of building their own rack-scale systems. If a hyperscaler accepts the argument that orchestration and data movement are the binding constraints, and that NVIDIA’s interconnect is the best solution, the rational move is to plug a custom ASIC into NVIDIA’s rack rather than build a competing full stack. That is the outcome NVIDIA wants. A hyperscaler that disagrees — and builds its own data-movement and interconnect layer around its own ASIC — is the scenario this argument is designed to preempt. The decision is not yet made, and it will not be made uniformly across all hyperscalers.
IMPLICATION #2 — For Reading the Jalapeno Benchmark Carefully
That OpenAI’s Jalapeno (co-developed with Broadcom) targets the same constraint Hardy names — minimizing data movement, reducing communication latency — is not a concession by NVIDIA; it is an independent data point surfaced by TechCrunch’s reporter. It suggests the underlying engineering constraint is real. But it also means OpenAI is designing its own solution to the same problem, not defaulting to NVIDIA’s. Both can be true simultaneously: the constraint is real, and the solutions are competing. See our Jalapeno benchmark deep-dive for the fine print on what those performance claims actually say.
IMPLICATION #3 — For Evaluating NVIDIA’s Claimed Numbers
The “upwards of 3x improvement in these operations” figure is NVIDIA’s own claim for a specific orchestration-offload operation on its own hardware, stated by an NVIDIA VP in an on-record interview. It is not an independent benchmark, not a third-party result, and not a system-wide performance number. Vera Rubin is currently rolling out. Until independent benchmarks exist at the system level — covering the full stack under realistic gigawatt-scale conditions — NVIDIA’s framing should be read as a
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
Sources: techcrunch.com · techcrunch.com · openai.com · nvidia.com · blogs.nvidia.com









