Ai2 releases Olmo-core 3 for open MoE training
The short version
Ai2 released Olmo-core 3, a redesigned open mixture-of-experts training stack combining expert parallelism, pipeline parallelism, and a distributed optimizer. The announcement includes trillion-parameter system tests and explicitly says random-routing benchmarks measure systems performance, not model quality.
What changed
The training stack moves from an earlier FSDP-based implementation to DDP, keeping experts resident on GPUs and routing data to them.
What it means for you
Researchers can use the open stack to study routing and parallelism; system benchmarks are not capability scores for a released model.
The announcement links a technical report, code, and an interactive parallelism walkthrough. Start with the experiment configurations and approaches the report did not adopt.
System capacity tests do not amount to a completed model release. Larger short-capacity tests also do not demonstrate sustained training performance. Event precision follows the article date, October 1.
Fact check
- VerifiedOlmo-core 3 is Ai2 open MoE training infrastructure combining expert parallelism, pipeline parallelism, and a distributed optimizer.Evidence
- VerifiedIt moves from the earlier FSDP-based stack to DDP with GPU-resident experts and routed data, reducing repeated weight gathering.Evidence
- VerifiedOfficial tests reach trillion-parameter scale; random routing measures system performance, not model quality. Researchers can use the open stack to study routing and parallelism.Evidence