Google’s Ironwood TPU Playbook: Optimizing Alibaba’s Qwen 3.5-397B as an Infrastructure Strategy

Based on Google’s Developers Blog engineering playbook, “Optimizing Qwen 3.5-397B MoE on Ironwood TPU7x.”

Google’s engineers have published a detailed optimization playbook for running Alibaba’s open-weights Qwen 3.5-397B on Ironwood TPUs — and the strategic signal it sends about where Google’s AI advantage actually sits is worth reading carefully.

Ironwood × Qwen 3.5-397B — Key Engineering Numbers (Apr–Jun 2026)

~3.1×

Throughput gain, decode-heavy workloads

~4.7×

Throughput gain, prefill-heavy workloads

~82%

Of Ironwood’s theoretical roofline reached (prefill)

192GB

Ironwood HBM vs ~288GB on Nvidia Blackwell GB300

What Happened

Google’s developers blog has published a systems-engineering playbook detailing how its teams optimized Alibaba’s Qwen 3.5-397B — a Chinese open-weights mixture-of-experts model with 397 billion parameters and 512 experts — to run on Google’s Ironwood (TPU v7x) accelerators. The work covers the period from April to June 2026. Google reports roughly 3.1× higher throughput on decode-heavy workloads and approximately 4.7× on prefill-heavy ones, reaching around 80–82% of Ironwood’s theoretical roofline — for example, 3,707 tokens/sec/chip against a 4,500 tokens/sec/chip prefill ceiling.

The engineering is substantive. The playbook describes a hybrid sharding scheme combining data parallelism and expert parallelism across the model’s 512 experts, custom JAX/Pallas kernels including a Ragged Page Attention rework reported to yield a 33.8% decode speedup, SparseCore-based MoE routing, and fused Grouped GEMM with FP8 precision. Critically, Google packaged the work to integrate with vLLM and SGLang — two of the most widely used open-source inference serving frameworks — positioning it as a drop-in path rather than a proprietary stack.

Important context before the strategy read: these are Google’s own self-reported figures, measured against Google’s own earlier TPU baseline from April — not a controlled head-to-head benchmark against Nvidia. Ironwood offers approximately 2,307 BF16 TFLOPS but only 192GB of HBM, versus roughly 288GB on Nvidia’s Blackwell GB300 — a real memory capacity gap the software engineering is partly working around. This is also strictly an inference (serving) story; Nvidia’s deepest moat remains in model training, a separate competition. And the post concludes by directing readers to Google Cloud’s TPU hub — this is developer marketing, and Ironwood is available only on Google Cloud, not as merchant silicon.

Context Timeline

April 2026

Google begins Qwen 3.5-397B optimization work on Ironwood TPU v7x; establishes baseline throughput numbers.

July 16, 2026

Gemini 3.5 Pro delay reported publicly, surfacing a frontier model gap for Google in competitive coding benchmarks.

July 17, 2026

Google publishes the Ironwood × Qwen 3.5-397B systems-engineering playbook, reporting ~3.1–4.7× throughput gains and ~80–82% roofline efficiency on inference workloads.

The Backdrop

Chinese open-weights models (Qwen, DeepSeek, Kimi, GLM) now represent a measurable and growing share of enterprise inference demand on open routing platforms.

The key insight: Google is not just optimizing a model — it is publishing a model-agnostic playbook. The choice to target Alibaba’s Qwen 3.5, integrate with vLLM and SGLang, and frame the work as reusable is a deliberate positioning move: Google Cloud as the efficient, open-standards home for whatever open-weights model enterprises actually want to run, regardless of who built it or where.

The Structural Read

Read the timing alongside yesterday’s Gemini 3.5 Pro delay and the picture sharpens considerably. On the same week Google acknowledged slipping at the frontier model layer, it published a detailed demonstration of TPU efficiency at the infrastructure layer. Whether coordinated or coincidental, the juxtaposition maps cleanly onto the AI value chain: when you are losing ground on the model, the rational play is to assert control over where models run.

The inference economy is the right frame here. As model weights proliferate — open-source, open-weights, multi-provider — the scarce and monetizable resource shifts from the model itself toward the compute substrate that serves it cheaply at scale. Reaching 80–82% of Ironwood’s roofline on a 397B MoE model, if the methodology holds, is a real engineering achievement. The strategic point is that Google is broadcasting that achievement as loudly as the engineering warrants.

The vLLM and SGLang integration is where the competitive logic becomes explicit. These are not Google frameworks — they are the frameworks that already run inference workloads on Nvidia hardware across thousands of enterprise deployments. By building Ironwood into that same ecosystem, Google is reducing the activation energy required to migrate off Nvidia for inference. This is a direct move on CUDA gravity at the serving layer — the same gravitational pull Apple could not escape even with its own silicon.

Map of AI — Infrastructure Layer

Google’s durable AI moat may be the TPU, not Gemini

In the Map of AI framework, the infrastructure layer — accelerator silicon, serving stacks, memory bandwidth — sits below the model layer and is structurally stickier. Models commoditize; efficient infrastructure compounds. A company that can serve any model, from any geography, at close to roofline efficiency on proprietary silicon it controls end-to-end occupies a durable position — especially if that silicon is only accessible through its own cloud. The playbook is genuine engineering; the business model underneath it is cloud lock-in wearing an open-standards jacket.

Google Developers Blog

“We demonstrate that it is possible to achieve high efficiency on Ironwood for a state-of-the-art open-weights MoE model, using a combination of hardware-aware sharding, custom kernels, and integration with popular open-source serving frameworks.”

Three Implications

IMPLICATION 1 — THE PLAYBOOK IS THE PRODUCT

Publishing a model-agnostic optimization guide with drop-in vLLM and SGLang support is Google reducing the switching cost off Nvidia for inference workloads. The target audience is not AI researchers — it is the enterprise platform team currently running Qwen, DeepSeek, or Llama on an Nvidia cluster and evaluating whether to move. The demand for Chinese open-weights models is already real and growing; Google is bidding to be where that demand runs.

IMPLICATION 2 — SHOWCASING A CHINESE OPEN MODEL IS A DELIBERATE SIGNAL

Choosing Alibaba’s Qwen 3.5 as the showcase model is not neutral. It positions Google Cloud as a geopolitically open, model-agnostic infrastructure layer at precisely the moment the chip market is bifurcating along US–China lines. Against the backdrop of Huawei gaining ground in Chinese AI infrastructure and export controls reshaping supply chains, Google is signaling that its cloud will run the models enterprises in any market actually want — monetizing the commoditization of model weights that Gemini is, at the moment, losing ground to.

IMPLICATION 3 — THE HEDGES ARE LOAD-BEARING

The 3.1–4.7× gains are measured against Google’s own April baseline, not against Nvidia hardware. Ironwood’s 192GB HBM versus Blackwell’s ~288GB is a real memory capacity constraint the custom kernels are engineering around, not eliminating. This is inference, not training — Nvidia’s strongest moat remains intact. And TPU access is Google Cloud-only, so the “open standards” framing coexists with genuine cloud lock-in. The engineering is real; the self-reported framing requires an independent benchmark before the claims can be taken at face value in procurement decisions.

Business Engineer Framework

The Map of AI — Infrastructure Layer

The Map of AI framework maps 200+ companies across nine layers of the AI stack. Google’s Ironwood playbook is a canonical infrastructure-layer move: asserting that the accelerator substrate — not the model — is where durable value accrues as weights commoditize. Understanding where each player sits in that stack, and which layers are compressing versus expanding in margin, is the core of the analysis.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Sources: developers.googleblog.com · developers.googleblog.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA