OpenAI and Broadcom’s Jalapeño Chip Posts Strong Inference Numbers — Believe the Direction, Scrutinize the Scoreboard

Based on OpenAI and Broadcom’s disclosures and reporting by TechSpot.

OpenAI has released its own benchmark numbers for Jalapeño, the custom inference ASIC it built with Broadcom — the results are directionally credible and structurally important, but they are self-reported on open-weight models, and the comparison is an inference-only chip against a general-purpose GPU.

Jalapeño — From Unveiling to Roadmap

Mid-2026

OpenAI and Broadcom publicly unveil Jalapeño, a custom inference ASIC fabricated at TSMC — designed from the ground up for LLM inference, with an approximately nine-month design-to-tape-out cycle.

August 2026

OpenAI reports Jalapeño’s self-run results on SemiAnalysis’s public InferenceX benchmark: 1.5–1.9× more work per watt at peak throughput and 1.7–3.6× lower end-to-end latency vs. Nvidia GB200/GB300. Rated at 700W; observed at or below 550W during test runs.

End-2026 (target)

Initial Jalapeño deployment — small volumes. Full production ramp expected 2027–2028. OpenAI targets approximately 50% lower inference cost per token versus GPU-based serving.

In Development

Generation 2 is deep in development; Generation 3 is taking shape. OpenAI has committed to a sustained, multi-generation silicon program — not a one-chip experiment.

What Happened

OpenAI and Broadcom have put detailed numbers behind Jalapeño — the custom inference ASIC the two companies unveiled earlier in 2026. According to OpenAI’s own disclosure, the chip was run on InferenceX, SemiAnalysis’s public, open-source inference benchmark — a fixed model set spanning GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — and the results were reported by OpenAI, not independently reproduced on the chip by a third party. That framing matters and belongs in the first sentence of any serious analysis of these numbers.

On that benchmark, OpenAI reports Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput, 1.7 to 3.6 times lower end-to-end latency, and 2.1 to 4.1 times higher performance on highly interactive workloads — all measured against Nvidia’s GB200 and GB300 systems. The chip is rated at 700 watts but ran at or below 550 watts during the tested runs, which is a favorable framing. The target is approximately 50% lower inference cost than GPU-based serving. End-2026 is the deployment target, at initial volumes; the real production ramp is a 2027–2028 story. Generation 2 is deep in development; Generation 3 is taking shape.

InferenceX is a credible, neutral framework — SemiAnalysis’s benchmark, not an OpenAI invention. But OpenAI ran Jalapeño on it and reported the results itself, the comparison is an inference-only ASIC against general-purpose GPUs, and the model set is open-weight, meaning these numbers say nothing about how Jalapeño performs on OpenAI’s own frontier models. None of that makes the numbers false. It makes them self-reported on a credible framework, which is a different confidence level than an independently audited head-to-head.

The key insight: An inference-only ASIC beating a general-purpose GPU on inference is the expected result — that is the entire logic of building an ASIC. The headline numbers are not the story. The cost savings, the independence from Nvidia’s pricing power, and the multi-generation roadmap behind this chip are the story. Believe the direction; keep a hand on the scoreboard.

Self-Reported Advantage vs. Nvidia GB200/GB300

OpenAI’s own InferenceX results — low end of reported ranges shown. Not independently reproduced. Inference-only ASIC vs. general-purpose GPU.

Work per watt at peak throughput 1.5×
Lower end-to-end latency 1.7×
Interactive workload performance 2.1×

Bars represent low-end multiples. Reported ranges: 1.5–1.9× (work/watt), 1.7–3.6× (latency), 2.1–4.1× (interactive). Source: OpenAI disclosure via TechSpot.

These are the multiples OpenAI reports for its Jalapeño inference chip against Nvidia's GB200 and GB300, shown
These are the multiples OpenAI reports for its Jalapeño inference chip against Nvidia’s GB200 and GB300, shown at the conservative low end of each range (the full ranges run to 1.9x on work per watt, 3.6x on lower latency, and 4.1x on interactive performance). Read them carefully: the benchmark is a public, open-source one from SemiAnalysis (InferenceX), but OpenAI ran Jalapeño on it and reported the results itself, comparing a purpose-built inference ASIC to general-purpose GPUs, on the benchmark’s open-weight models rather than OpenAI’s frontier ones – and the chip is not yet in production. An ASIC beating a GPU on the one job it was designed for is the expected outcome, not a surprise. Source: OpenAI, on the InferenceX benchmark (self-reported).

The Structural Read

Start with the category. An ASIC is a chip built to do one thing; a GPU is built to do many. When you design straight at the memory bottleneck that limits GPU efficiency on LLM inference — and strip out everything the chip does not need for that one job — you are supposed to win on performance per watt for that job. This is precisely the story of Google’s TPU, now a decade old. The 1.5-to-1.9× per-watt figure, impressive as it is, is a feature of the category, not a miracle of engineering. Treating it as a shock requires forgetting what ASICs are for. The Beyond Nvidia’s Moat analysis at Business Engineer sets this up precisely: the ASIC-versus-GPU comparison on inference is apples-to-oranges in Jalapeño’s favor on this one axis, while Nvidia’s chips also train, flex across workloads, and carry a decade of software ecosystem that an ASIC does not inherit.

The benchmark framing compounds this. InferenceX is SemiAnalysis’s public, open-source benchmark — its model set is fixed (GPT-OSS 120B, DeepSeek R1 670B, Kimi K2.5 1T), and OpenAI did not choose those models. But OpenAI ran Jalapeño on that benchmark and reported the results itself, comparing against the best available GB200/GB300 numbers, without an independent party reproducing the Jalapeño run. That is self-reported on a neutral framework — directionally credible, not independently verified. The open-weight model set is also a meaningful gap: it says nothing about how Jalapeño performs on o3, o4-mini, or GPT-5 — the workloads that actually drive OpenAI’s revenue and cost structure. And the chip is not deployed yet. End-2026 is a target for initial volumes; production at scale, where power delivery, yield, networking, and software maturity all bite, is a 2027–2028 story.

None of those caveats make the numbers false, and they matter less than they might — because beating Nvidia on a benchmark was never the actual prize. OpenAI burns billions of dollars a quarter, and inference is its single largest cost. As The AI Value Chain framework makes clear, the lab that controls its own compute layer captures a structurally different margin profile than one that rents GPUs from Nvidia at whatever price Nvidia sets. Nvidia has been raising GPU prices and memory remains the binding constraint — a chip that serves OpenAI’s models at roughly half the cost transforms its unit economics whether or not every headline multiple holds under independent scrutiny.

The Framework

The Multi-Generation Roadmap Is the Real Signal

Gen 1 deploying end-2026, Gen 2 deep in development, Gen 3 taking shape — this is not a chip, it is a silicon program. OpenAI is making the same vertical-integration commitment that Anthropic is making with its TPU strategy and that every major lab is pursuing through infrastructure vertical integration: own the metal your models run on, because whoever controls the compute layer sets the economics for everyone above it. The benchmark numbers are a snapshot; the roadmap is a structural commitment.

Three Implications

IMPLICATION 1 — NVIDIA’S INFERENCE MARGINS FACE STRUCTURAL PRESSURE

Jalapeño does not render Nvidia irrelevant — training, flexibility, and software moat remain intact — but it begins to erode the inference rent. If the approximately 50% cost target holds at production scale, and if Gen 2 and Gen 3 improve on it, OpenAI progressively reduces the portion of its cost structure that flows to Nvidia. That is a direct challenge to Nvidia’s inference pricing power, not its existence. The Beyond Nvidia’s Moat analysis frames exactly this dynamic: the moat is real, but inference is the layer most vulnerable to custom silicon.

IMPLICATION 2 — UNIT ECONOMICS, NOT BENCHMARK POSITIONS, ARE THE PRIZE

The scoreboard headline — 1.5 to 4.1× depending on workload and metric — is less important than the cost-per-token trajectory. OpenAI’s business model depends on serving an enormous and growing inference load profitably. A chip that halves that cost, even roughly, changes the margin structure of every product it runs: ChatGPT, the API, operator deployments. The exact multiples matter less than the direction; the direction is almost certainly real even if every self-reported number is generous.

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA