Allen AI released Olmo-core 3 on October 1, an open-source training framework for mixture-of-experts models that can scale past one trillion parameters without the throughput collapse that typically accompanies larger expert pools.
By switching the communication strategy to keep expert weights resident on GPUs and routing data to them, the stack cuts the overhead that has made MoE training prohibitively expensive for many academic groups and smaller labs, according to a Hugging Face blog post from the research group.
The framework builds on Allen AI’s earlier MoE work in OlmoE, which used 64 routed experts, and its dense Olmo 3 model. The previous Olmo-core MoE implementation relied on fully sharded data parallelism (FSDP), which gathered and reshared weights for every batch. Olmo-core 3 replaces that with a distributed data parallelism (DDP) approach, avoiding repeated weight gathering and improving throughput. In a test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU—2.7 times the 19,400 tokens per second achieved under the old FSDP system, the team reported.
The system combines expert parallelism, pipeline parallelism, and a distributed optimizer to split model, layers, and optimizer state across GPUs. Routing optimizations include rowwise expert parallelism to place data directly into expert input buffers and GPU-resident routing that keeps metadata on the GPU to avoid CPU wait times. Grouped GEMM combines small expert computations for higher GPU utilization.
On four B300 GPUs, enabling MXFP8 mixed precision across feed-forward and expert-data movement operations raised training throughput 21% over a BF16 baseline while reducing peak active memory from 103 GiB to 95 GiB, the post states.
Crucially, the team reported that expanding the expert pool from 8 to 128 experts—while still selecting only four per token—increased total parameter count from 4.6 billion to 47 billion with a throughput drop of less than 5%. The stack was benchmarked on a 1.2-trillion-parameter configuration using 512 GPUs, achieving 858 TFLOP/s/GPU with random routing. A short-capacity test using DeepEP v2 reached a 2.38-trillion-parameter setup, though sustained training performance was not provided.
The technical report also documents findings with practical implications: a load-balancing score could improve while actual workload balance declined—a phenomenon the authors call token gerrymandering—and lowering expert learning rates for less-frequently used experts did not help. Overlapping communication and computation on separate GPU streams sometimes slowed end-to-end training.
Allen AI did not release independent test results or external benchmarks alongside the code. The announcement says the next generation of Olmo will use an MoE architecture, aiming to be the lab’s most capable model yet, trained on its largest dataset and longest context window. The open-source code and technical report are available on GitHub and via the blog.