Xiaomi Published What Its Training Run Cost — and That Changes the Comparison

Xiaomi released weights, training tools, and a partial cost figure for two frontier models — and that last disclosure reframes what the leaderboard is actually measuring.

What Happened

Xiaomi published MiMo-V2.6 on 21 September 2026, according to VentureBeat. The release covers two models — Pro, at approximately 1.02 trillion parameters, and Flash, a sparse mixture-of-experts model with 309 billion total parameters and roughly 15 billion active — along with weights, code, and a technical report, all under an MIT licence.

Alongside the models, Xiaomi released more than 7,000 reinforcement-learning environments and training tools under the same licence. Then it published something the vast majority of model releases omit: what a defined phase of training cost. The Pro reinforcement-learning run completed 30 steps across approximately 750,000 trajectories in under six days for approximately $2.62 million. The Flash RL run completed the same pipeline for approximately $850,000. Both figures cover the reinforcement-learning phase only. No pretraining cost has been disclosed for either model, and that distinction matters for every comparison those numbers might enter.

On the third-party Artificial Analysis Intelligence Index — the benchmark carrying the most weight here precisely because it is independently constructed — Pro scored 46, tying Grok 4.7 and taking the top open-weights position. Additional benchmark results reached us through secondary reporting and are not independently verified here; those figures are addressed in the structural section below, along with what they actually show.

The key insight: Publishing a cost figure — even a partial one — does not just add data to a leaderboard. It changes the question the leaderboard answers. Once a denominator exists, the competition quietly shifts from “which model is better” to “how much capability per unit of spend.” That is a different contest, with a different set of likely winners, and Xiaomi introduced it unilaterally.

Top of open weights on one third-party index; second here. Both are true, and that is the point.
Top of open weights on one third-party index; second here. Both are true, and that is the point.

The Benchmark Picture — Mixed, as Reported

The honest read of the scoreboard is a mixed one, and a mixed result is more useful than a clean narrative. The figures below (except the Artificial Analysis index) come via secondary reporting and are not independently verified here.

Artificial Analysis Intelligence Index (third-party, independently verified)

MiMo-V2.6 Pro / Grok 4.7 (tied) 46

AutomationBench — as reported; not independently verified here

DeepSeek V4.1 Flash 54.8
MiMo-V2.6 Pro 53.1
GPT-6 Astra 52.0
Claude Opus 5 50.3
Claude Fable 5 46.2

DeepSWE v1.1 — as reported; not independently verified here

DeepSeek V4.1 Flash / Claude Opus 5 (tied ~74) ~74
MiMo-V2.6 Pro 71.9
Claude Fable 5 70.0

On AutomationBench and DeepSWE v1.1, as reported through secondary sources and not verified here, Pro sits below DeepSeek V4.1 Flash on both, and below Claude Opus 5 on DeepSWE. Leading one index while trailing others is not a contradiction — different benchmarks measure different work, and any composite index embeds the weighting choices of whoever built it. Nothing here says Pro beats, matches, or overtakes closed models in general.

The Structural Read

Three structural moves sit inside this release, and they operate at different layers of the stack. The cost disclosure is the most visible. The environments release is arguably the most consequential. And the benchmark result itself — mixed but real — is the one most likely to be reported inaccurately.

On the cost figure: Xiaomi has not stated why it published the RL-phase cost, and nothing here claims to know its reasoning. What follows is inference from structure, not attributed motive. A leaderboard that shows only scores has one available question: which model performs better. Adding a cost figure — even a partial one — installs a denominator and produces a second question: how much capability per unit of spend. That second question is a different competition. A participant who expects to look strong on that ratio has a structural interest in the ratio existing at all. Once one participant publishes a denominator, silence from others is no longer neutral — it becomes legible, because everyone can see that disclosure was possible.

One critical precision must travel with both numbers: the $2.62 million and $850,000 figures cover the reinforcement-learning phase only and are not total training costs. No pretraining cost has been disclosed. Setting these figures alongside another company’s total training cost would compare a disclosed component against an undisclosed whole — a counting error that produces a conclusion neither figure actually supports.

Inference From Structure — Not Stated Reasoning

“Releasing weights lets other people use a model. Releasing more than 7,000 RL environments and training tools lets other people produce one — including models the original author never anticipated and would not necessarily welcome. The second kind of release compounds; its beneficiaries can improve on what they were given and return that improvement to the same commons. A set of weights is a fixed artefact that ages. An environment is a means of production.”

The environments release is the part receiving least attention and may be the structurally largest. This is what “open” can mean at different layers of a stack, and the layer matters more than the label. No adoption figure for the environments appears here, and nothing here claims anyone has used them or will.

Map of AI — Where This Release Lands

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

The $2.62 million and $850,000 figures cover the reinforcement-learning phase only — they are not the total cost of training either model, and no pretraining cost has been disclosed. This is not investment advice. Setting either number beside another company’s total training cost would compare a part with a whole. The Artificial Analysis score is a third-party measurement; the remaining benchmark figures reached this piece through secondary reporting and are not independently verified here. The results are mixed: MiMo-V2.6-Pro tops the open-weights slot on that index while sitting below DeepSeek V4.1 Flash on AutomationBench and below both DeepSeek V4.1 Flash and Claude Opus 5 on DeepSWE v1.1 — nothing above claims it beats, matches or overtakes closed models in general. Xiaomi has not stated why it published these figures, and every strategic reading above is inference from structure rather than the company’s reasoning. No compute, chip, inference-price, adoption or revenue figure appears, nothing above claims anyone has used the released RL environments, and no position is taken on export controls, trade policy or any government.

Sources: venturebeat.com · aiweekly.co · orcarouter.ai

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA