Ai2’s Olmo-core 3: 2.7x MoE Training Throughput, Open Stack

Every claim here is Ai2’s own statement in its blog post about its own system. This publication read the post text only and verified none of the results independently.

Ai2 released Olmo-core 3 on 1 October 2026, an open training framework with a redesigned system for mixture-of-experts models. In a preliminary test on eight NVIDIA B300 GPUs, Ai2 says a 47-billion-parameter MoE processed 52,000 tokens per second per GPU, against 19,400 with its earlier implementation, about 2.7 times the throughput. Ai2 says its trillion-parameter tests used random routing and measure system performance, not model quality. This publication read Ai2’s own post and verified none of its results independently.

What Ai2 Released

Ai2 describes Olmo-core 3 as “a significant upgrade to our framework for developing large language models.” It says the system is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency, and that it is one of the core systems behind the next generation of Olmo.

Ai2 says the framework is fully open, so researchers and developers can use it to train their own MoEs, adapt it to different hardware and experiment with routing and parallelism. It says Olmo 3 used a dense architecture, while its earlier sparse work, OlmoE, used 64 routed experts.

From Ai2's post: in a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000
From Ai2’s post: in a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with Olmo-core 3, against 19,400 with Ai2’s earlier implementation, which Ai2 calls about 2.7 times the throughput.

Business Pill · MANY SPECIALISTS, FEW AT WORK

A one-minute explainer of the idea behind this story: mixture of experts. It teaches the concept, not this story’s figures.

The key insight: Ai2’s headline number is a before-and-after on its own code, not a comparison with another lab’s system. The 2.7 times comes from a preliminary eight-GPU test, and the trillion-parameter runs measure system performance with random routing, not the quality of a trained model.

The Expert-Pool Benchmark

In one benchmark, Ai2 says it increased the expert pool from 8 to 128 while still selecting four experts per token, keeping the active parameters per token at about 3.2 billion. Total parameter capacity grew from 4.6 billion to 47 billion, and Ai2 says training throughput fell by less than 5%.

Ai2 says the same infrastructure has been benchmarked at over one trillion total parameters.

The 2.7x Test

Ai2 says its earlier MoE implementation used fully sharded data parallelism, which gathers and reshards model weights for each small batch. Olmo-core 3 switches to a system based on distributed data parallelism that keeps experts resident on GPUs and routes the data to them.

In what Ai2 calls a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the earlier implementation. Ai2 calls that about 2.7 times the throughput.

Lower Precision and Routing

Ai2 says Olmo-core 3 supports MXFP8, a lower-precision number format. In a controlled benchmark on four B300 GPUs, with work distributed uniformly across experts and MXFP8 enabled across the parts of the system where it helped most, it says throughput was about 21% higher than with BF16, and peak active memory fell from 103 GiB to 95 GiB.

Ai2 also lists changes that reduce the cost of routing data to the experts and running their computations: rowwise expert parallelism, GPU-resident routing and grouped GEMM. It says these techniques “have to work together,” because speeding up one part of training can create costs elsewhere.

Trillion-Parameter Tests and Their Limits

Ai2 says it benchmarked a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 B300 GPUs, with a highest observed throughput of 858 TFLOP/s per GPU. It says these tests used random routing to measure system performance, rather than the quality of a trained model.

Ai2 also says it reached a configuration with 2.38 trillion total parameters using an alternative communication method, DeepEP v2. It calls this a short-capacity test rather than a full training run, which shows the scale Olmo-core 3 can reach rather than sustained training performance.

Findings From Its Experiments

Ai2 says its technical report documents experiments that informed how it trains MoEs. It says a score meant to encourage balanced routing could improve even as the actual workload became less balanced, which it calls token gerrymandering.

It also says lowering experts’ learning rates because they process fewer tokens did not improve results in the model family it tested. Overlapping communication and computation did not always make training faster, and in some tests it slowed end-to-end execution.

The Structural Read

The post makes three kinds of claim at three levels of evidence. The 2.7 times comes from a preliminary test on eight GPUs against Ai2’s earlier implementation. The 21% from MXFP8 comes from a controlled benchmark on four GPUs with work spread uniformly across experts. The 1.2-trillion-parameter and 2.38-trillion-parameter results are capacity demonstrations: random routing for the first, a short-capacity test for the second.

The expert-pool benchmark is a statement about cost per token. Ai2 says it grew total capacity from 4.6 billion to 47 billion parameters while keeping about 3.2 billion active per token, and that training throughput fell by less than 5%. The post reports no model-quality result for these runs.

Ai2 calls NVIDIA’s Megatron-Core an established option for training large MoEs, and says Olmo-core 3 improves throughput over Ai2’s own earlier FSDP-based implementation. The post gives no head-to-head figure against Megatron-Core.

Ai2 — Olmo-core 3, 1 October 2026

“These tests used random routing to measure system performance, rather than the quality of a trained model.”

Three Implications

OPEN INFRASTRUCTURE, NOT AN OPEN MODEL What Ai2 released is the training stack. It says its next-generation Olmo will use an MoE architecture and that it is aiming for it to be its most capable Olmo yet; the post gives no release date. In Ai2’s words, “model weights are more useful when the infrastructure and training decisions behind them are open too.”

COMPUTE EFFICIENCY IS THE PITCH Ai2 opens by saying training large models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. Every figure in the post is a throughput, memory or capacity figure, and none is a dollar cost.

WHAT REMAINS UNKNOWN The post reports no trained model from the stack, no head-to-head with Megatron-Core, no release date for the next Olmo and no independent replication. This publication did not read the technical report or run the code.

What Is Not Established

Ai2 says its next-generation Olmo will use an MoE architecture and that it is aiming for it to be its most capable Olmo yet. The text of the post gives no release date and reports no trained model from this stack.

This publication read the post text only. It did not read the technical report, run the code or try the interactive demo, and did not seek a response from Ai2.

Business Engineer Framework

The Map of AI — Where Training Infrastructure Sits in the Stack

The Map of AI places more than 200 companies across nine layers of the stack, from silicon to application. Open training software like this sits between the chips and the models, and the Map shows who occupies the layers around it.

Read the Map of AI →

The Bottom Line

Ai2 says its open Olmo-core 3 stack processed 52,000 tokens per second per GPU on a 47-billion-parameter MoE in a preliminary eight-GPU test, against 19,400 for its earlier implementation. It says it benchmarked a 1.2-trillion-parameter model across 512 B300 GPUs using random routing, which measures system performance and not model quality. This publication verified none of it independently.

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

This piece rests on Ai2’s blog post of 1 October 2026. This publication read the post text only, did not read the technical report, run the code or try the interactive demo, and did not seek a response from Ai2. Nothing above predicts anything, and nothing here is investment advice.

Sources: allenai.org · Ai2 blog post ‘Olmo-core 3: Open, scalable training infrastructure for large MoEs’, 1 October 2026 (post text as supplied)

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA