Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

2026-10-02 · Hugging Face

Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Large MoEs

Overview

Ai2 has released Olmo-core 3, a significant upgrade to its framework for developing large language models, featuring a redesigned open mixture-of-experts (MoE) training system. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It is one of the core systems behind the next generation of Olmo and part of Ai2's ongoing commitment to open up the tools and training infrastructure behind each new model.

Challenges of MoE Training

Training large AI models requires substantial compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach—they can contain many more parameters without requiring every input to use all of them. However, the full model still must be stored across GPU memory and updated during training, and directing inputs to the right experts across a cluster creates additional communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model per input.

In one benchmark, Olmo-core 3 increased the expert pool from 8 to 128 while still selecting only four experts per token, keeping active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%. The same infrastructure has been benchmarked at over one trillion total parameters.

Building a Training Stack Around How MoEs Work

Olmo-core has evolved with each generation of Olmo. Ai2's work on sparse models traces back to OlmoE, which used an MoE architecture with 64 routed experts. Olmo 3, by contrast, used a dense architecture where nearly all of the model was active for every token. Olmo-core 3 extends the framework with a training system designed for much larger MoE models.

The earlier MoE implementation in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism (DDP), keeping experts resident on GPUs and routing relevant data to them, avoiding repeated weight gathering.

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the earlier implementation—approximately 2.7× the throughput.

Scaling and Optimizing MoE Training

Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations for more efficient routing and computation.

Three Distribution Techniques

  • Expert parallelism spreads experts across GPUs, so each GPU stores only part of the full expert pool.
  • Pipeline parallelism splits the model's layers across groups of GPUs, reducing how much of the model each GPU needs to keep in memory.
  • Distributed optimizer spreads the optimizer state across GPUs instead of storing a full copy on every GPU.

Routing and Computation Optimizations

  • Rowwise expert parallelism places routed data directly into expert input buffers, minimizing rearrangement overhead.
  • GPU-resident routing keeps routing metadata on GPUs, so the CPU can queue work without waiting for information to be copied back.
  • Grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.

MXFP8 Support

Olmo-core 3 supports MXFP8, a lower-precision number format that represents some values with fewer bits, reducing computation and data movement between GPUs. In a controlled benchmark on four NVIDIA B300 GPUs, enabling MXFP8 yielded approximately 21% higher training throughput than the BF16 baseline, while peak active memory fell from 103 GiB to 95 GiB. Most gains came from feed-forward computation and moving data between experts rather than attention alone.

These techniques and optimizations must work together—speeding up one part of training can create costs elsewhere. Olmo-core 3 is built around those trade-offs across the full training process, giving researchers control over how the pieces fit together.

Source