Trending: On-device modelsSearch
iHeartGeek
iTECH

Ai2's Olmo-core 3 brings open MoE training to trillion scale

Ai2 has rebuilt the training framework behind Olmo around a new mixture-of-experts system that holds throughput steady as models grow into the trillion-parameter range.

The Olmo-core 3 release title card from Ai2, showing the project name in pink on a teal background.

Ai2 has released Olmo-core 3, a rework of the framework it uses to train the Olmo family of language models. The centrepiece is a redesigned mixture-of-experts training system, built to take MoE models into the trillion-parameter range without giving back the efficiency that makes the architecture worth using.

The problem is communication. An MoE model is cheap to run because each input touches only a handful of specialists, but the whole model still has to sit in GPU memory, and routing data to the right expert across a cluster burns time and bandwidth. As the expert pool grows, those costs can swallow the savings that come from activating only part of the model.

The benchmark numbers

In one test, Ai2 grew the expert pool from 8 to 128 while still selecting four experts per token, holding active parameters at roughly 3.2 billion. Total capacity climbed from 4.6 billion parameters to 47 billion, and training throughput fell by less than 5%. The same stack has been benchmarked at more than a trillion total parameters.

The new implementation drops the fully sharded data parallelism of the earlier Olmo-core MoE work in favour of distributed data parallelism. Experts stay resident on their GPUs and the data is routed to them, which removes the repeated gathering and resharding of weights. On eight Nvidia B300 GPUs, a 47-billion-parameter MoE pushed 52,000 tokens per second per GPU, against 19,400 on the older stack - about 2.7 times the throughput.

What else is new

Olmo-core 3 spreads a model across hardware with expert parallelism, pipeline parallelism and a distributed optimizer, and trims routing overhead with rowwise expert parallelism, GPU-resident routing and grouped GEMM. Support for MXFP8, a lower-precision number format, lifted end-to-end throughput by about 21% over a BF16 baseline in a four-B300 test, while peak active memory dropped from 103GiB to 95GiB.

The largest configuration Ai2 reports is a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs, peaking at 858 TFLOP/s per GPU under random routing. A separate test using DeepEP v2 reached 2.38 trillion total parameters, though Ai2 describes that as a short-capacity check rather than a sustained training run.

Our opinion

Publishing the weights is the easy half of open model development; publishing the machinery that produced them is the half that lets other labs actually catch up. Olmo-core 3 is a system for training mixture-of-experts models at a scale that used to demand a hyperscaler's internal stack, and the valuable part is not the headline number but the write-up of what did not work - the token gerrymandering failure, the overlapping GPU streams that slowed training down. Those dead ends are usually the first thing a company with a moat keeps to itself.