Theo Browne’s read on GLM-5.3-Flash is directionally right but causally backwards — the non-Nvidia chips prove sovereignty, not cheapness.
A model called “Ox Alpha” quietly topped OpenRouter usage charts before being unmasked on August 27–28, 2026 as GLM-5.3-Flash from Z.ai — the commercial arm of Zhipu AI, a Chinese lab. Theo Browne caught the story early. His instinct about the chip independence was correct. His causal chain was not.
Z.ai’s stated claim: the entire preview — roughly 100 trillion tokens per day — ran entirely on domestically-developed Chinese chips, no Nvidia in the serving stack. That is a genuinely remarkable claim. The causation problem is what Theo attached to it.
The key insight: Z.ai’s own position is cost parity with Nvidia, not a cost advantage. Domestic chips are not cheaper per FLOP. The cheapness comes from the architecture — approximately 18B active parameters in a sparse mixture-of-experts design out of roughly 320B total, priced at about $0.15 per million input tokens and $0.50 per million output tokens (roughly a tenth of GLM-5.3). Few FLOPs per token is the price lever. The silicon is the sovereignty lever. Theo inverted the two.
The Structural Read
The strategic significance of Ox Alpha is not the price. It is the decoupling. If a Chinese lab can serve a top-usage model at 100 trillion tokens per day entirely on domestic silicon, export controls on Nvidia stop being a choke point for inference. That is a compute-sovereignty proof point, not a cost story.
The barbell is important here. Cheap open Chinese models on sovereign chips commoditize the inference floor. The frontier training race remains heavily Nvidia-bound. Serving at inference scale on domestic chips is not the same problem as training at frontier scale — the harder problem is still largely unsolved without Nvidia-class hardware. These are two different races.
The honest counter deserves prominent space: Z.ai did not name its chip supplier. “Huawei Ascend” is analyst inference — GLM-5.1 and 5.2 were trained on Ascend hardware, and Cambricon is another named candidate — but it is not company-confirmed. The 100 trillion tokens per day figure and the all-domestic-silicon claim are Z.ai’s own stated numbers and have not been independently audited. Treat the hardware claim accordingly.
CHEAPNESS IS THE ARCHITECTURE
Roughly 18B active parameters in a sparse MoE design means very few FLOPs per token. That is the price mechanism. At approximately $0.15 per million input tokens, GLM-5.3-Flash competes on cost through design, not through cheaper silicon.
SOVEREIGNTY IS THE SILICON
A top-usage model served at claimed scale with zero Nvidia in the stack — supplier unnamed, claim unaudited — is a proof-of-concept for inference independence. If it holds up, export controls lose their leverage at the serving layer.
THE HONEST LIMIT
The hardware claim is company-stated and unaudited. The chip supplier is unnamed. Parity with Nvidia is the stated position, not an advantage. Do not read this as “China doesn’t need Nvidia” — the frontier training problem is a different and harder constraint entirely.
The Bottom Line
The cheap price is the architecture. The non-Nvidia chips are the sovereignty story. Conflating them misses both. The real headline from Ox Alpha is not “cheap because non-Nvidia” — it is that a top model was served at massive claimed scale with zero Nvidia, at stated cost parity, on chips the company will not name. That is bigger than the price point, and more honest than the shortcut.
Clip via Theo (t3.gg) (source).





