The Allen Institute for AI (Ai2) has released Olmo-core 3, an open-source training framework designed to support large Mixture-of-Experts (MoE) models. The release introduces a redesigned training stack intended to scale MoE architectures into the trillion-parameter range while maintaining computational efficiency.
What Happened
Olmo-core 3 replaces the fully sharded data parallelism (FSDP) used in previous versions with a system based on distributed data parallelism (DDP). This change keeps experts resident on GPUs and routes data to them, avoiding the repeated weight gathering required by the earlier implementation. According to Ai2, in preliminary tests on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the previous FSDP-based implementation, representing approximately 2.7x the throughput.
The framework integrates several techniques to distribute models across hardware, including expert parallelism, pipeline parallelism, and a distributed optimizer. It also incorporates optimizations such as rowwise expert parallelism, GPU-resident routing, and grouped GEMM. Additionally, Olmo-core 3 supports MXFP8, a lower-precision number format. In a controlled benchmark on four NVIDIA B300 GPUs, enabling MXFP8 resulted in about 21% higher training throughput compared to the BF16 baseline, while peak active memory fell from 103 GiB to 95 GiB.
Ai2 reports that the infrastructure has been benchmarked on a 1.2-trillion-parameter model across 512 GPUs, achieving a highest observed throughput of 858 TFLOP/s/GPU. The team also experimented with DeepEP v2 to reach configurations with 2.38 trillion total parameters, though this was a short-capacity test rather than a full training run.
Why It Matters
The release addresses the growing computational costs and energy demands of training large AI models. While MoE architectures offer efficiency by activating only parts of the model per input, coordinating these experts across GPU clusters creates significant communication overhead. Olmo-core 3 aims to mitigate these costs, potentially making advanced model development more accessible to academic researchers and smaller labs that lack the resources of frontier labs.
The framework is fully open, allowing researchers to adapt it to different hardware and experiment with routing and parallelism strategies. Ai2 positions this release as part of a broader commitment to open model development, arguing that model weights are more useful when the underlying infrastructure and training decisions are also transparent. The new stack serves as the foundation for the next generation of Olmo models, which will use an MoE architecture.
The Bottom Line
Olmo-core 3 provides an open, scalable training infrastructure for large MoE models, offering significant throughput improvements over previous implementations and supporting scaling into the trillion-parameter range. The release emphasizes transparency and accessibility, providing tools and technical reports that detail the system design and experimental findings for the broader research community.