Aivora
Hugging FaceLLMAdvanced

Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoEs

Olmo-core 3 登場:開源兆級參數 MoE 模型訓練基礎架構

2 min read
Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Trillion-Parameter MoEs
The 30-second version

Olmo-core 3 redesigns the MoE training stack, transitioning from FSDP to a DDP-based system that keeps experts resident on GPUs. In benchmarks using NVIDIA B300 GPUs, a 47B MoE model achieved 52k tokens/sec per GPU, representing a 2.7x throughput increase over previous versions. Supporting MXFP8 precision and advanced parallelism, Olmo-core 3 scales up to trillion-parameter configurations, offering researchers an efficient, fully open-source infrastructure.

Key points

01

GPU-Resident Experts

Switches from FSDP to DDP, keeping experts resident on GPUs and routing data to them, eliminating repeated weight-gathering overhead.

02

2.7x Throughput Boost

In a 47B MoE benchmark on 8 B300 GPUs, throughput increased from 19.4k to 52k tokens per second per GPU.

03

Trillion-Parameter Scale

Successfully benchmarked a 1.2-trillion-parameter model across 512 GPUs, and reached 2.38 trillion parameters in a capacity test using DeepEP v2.

04

MXFP8 Precision Support

Enabling MXFP8 precision boosts training throughput by roughly 21% compared to BF16, while reducing peak memory from 103 GiB to 95 GiB.

How it works

Comparison of Olmo-core Implementations
舊版 Olmo-core (FSDP)全新 Olmo-core 3 (DDP)
Expert Residency頻繁重新聚集與切分權重 (Gather & Reshard)常駐於 GPU 並直接路由資料 (Resident on GPUs)
47B MoE Throughput19,400 tokens/sec/GPU52,000 tokens/sec/GPU (提升 2.7 倍)
Max Parameter Scale數十億參數級別 (Billions)最高達 2.38 兆參數 (Trillion-scale)

Why it matters

Training large MoE models is historically costly due to communication bottlenecks, locking academic researchers and small labs out. By open-sourcing Olmo-core 3, the Allen Institute for AI lowers the barrier to training trillion-parameter models. This framework provides open, efficient infrastructure alongside the model weights, fostering transparent and collaborative AI development.

Who it affects

  • AI Developer
  • AI Researcher
  • Enterprise Leader

How to use it

  1. 1Academic institutions and open-source communities training ultra-large MoE models on existing GPU clusters.
  2. 2Leveraging MXFP8 low-precision and distributed optimizers to reduce hardware memory footprints during large-scale training.

Limitations & caveats

  • Overlapping communication and computation on separate GPU streams did not always make training faster, sometimes slowing execution.
  • Subject to 'token gerrymandering' where balanced routing scores improve even as the actual workload becomes less balanced.

Related

SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
arXivLLM

SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability

SCAPO:藉由半事實穩定性最佳化 RLVR 的 Token 級信用分配

SCAPO is a novel variant of GRPO that incorporates semifactual stability into token-level credit assignment, significantly improving LLM reasoning accuracy and out-of-distribution generalization.

2 min read