Product update

Ai2 releases Olmo-core 3 for open MoE training

Event time 1 independent sourceEditorial score 67/100Updated here

The short version

Ai2 released Olmo-core 3, a redesigned open mixture-of-experts training stack combining expert parallelism, pipeline parallelism, and a distributed optimizer. The announcement includes trillion-parameter system tests and explicitly says random-routing benchmarks measure systems performance, not model quality.

What changed

The training stack moves from an earlier FSDP-based implementation to DDP, keeping experts resident on GPUs and routing data to them.

What it means for you

Researchers can use the open stack to study routing and parallelism; system benchmarks are not capability scores for a released model.

The announcement links a technical report, code, and an interactive parallelism walkthrough. Start with the experiment configurations and approaches the report did not adopt.

System capacity tests do not amount to a completed model release. Larger short-capacity tests also do not demonstrate sustained training performance. Event precision follows the article date, October 1.

Fact check

  • VerifiedOlmo-core 3 is Ai2 open MoE training infrastructure combining expert parallelism, pipeline parallelism, and a distributed optimizer.Evidence
  • VerifiedIt moves from the earlier FSDP-based stack to DDP with GPU-resident experts and routed data, reducing repeated weight gathering.Evidence
  • VerifiedOfficial tests reach trillion-parameter scale; random routing measures system performance, not model quality. Researchers can use the open stack to study routing and parallelism.Evidence

Coverage timeline

  1. Primary sourceAi2
    Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs