Scaling Laws for Looped Mixture of Experts: Combining Recurrence and Sparsity
融合循環與稀疏架構:Looped MoE 的全新縮放定律
This paper presents the Loop Scaling Laws, the first framework to jointly model recurrence and sparsity alongside model size and data. While recurrence increases computational depth without adding parameters, MoE expands model capacity without increasing active compute. The study shows that sparsity yields ~3x active-parameter efficiency, and recurrence delivers ~2x total-parameter efficiency on reasoning tasks. At a trillion-token scale, a Looped MoE matches the performance of a non-looped MoE twice its size, while enabling dynamic test-time scaling.
Key points
First Joint Scaling Law
Successfully unifies recurrence and MoE sparsity into a single mathematical framework, predicting held-out loss more accurately than prior methods.
Complementary Efficiency Gains
Sparsity delivers ~3x active-parameter efficiency, while recurrence yields ~2x total-parameter efficiency on complex reasoning tasks.
Trillion-Token Validation
At a trillion-token scale, a Looped MoE matches a ~2x larger non-looped model under matched compute, enabling dynamic test-time scaling.
How it works
| Dense (稠密模型) | Standard MoE (標準 MoE) | Looped MoE (迴圈 MoE) | |
|---|---|---|---|
| Active Param Efficiency | 基準 (1x) | 高 (~3x 效率) | 高 (~3x 效率) |
| Total Param Efficiency | 基準 (1x) | 基準 (1x) | 極高 (~2x 推理效率) |
| Test-Time Scaling | 不支援 | 不支援 | 支援 (調整遞迴次數) |
Why it matters
This research breaks the bottleneck of traditional model scaling by showing that recurrence and MoE are highly synergistic. It provides a principled design foundation under tight memory and compute constraints. Developers can now build models with smaller memory footprints that still leverage the capacity of MoE. The ability to perform test-time scaling via variable recurrence steps opens up new avenues for efficient inference and adaptable deployment.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1Resource-Constrained Deployment: Running high-performance LLMs on edge devices or consumer hardware with strict memory and compute limits.
- 2Test-Time Dynamic Scaling: Dynamically adjusting recurrence steps during inference based on task complexity to optimize operational compute costs.
Limitations & caveats
- Recurrent architectures may increase sequential inference latency due to the serial nature of deep loops, limiting parallelization.
- The scaling law parameters may require re-calibration when applied to highly heterogeneous datasets or novel MoE routing mechanisms.
Related
SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
SCAPO:藉由半事實穩定性最佳化 RLVR 的 Token 級信用分配
SCAPO is a novel variant of GRPO that incorporates semifactual stability into token-level credit assignment, significantly improving LLM reasoning accuracy and out-of-distribution generalization.
STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸
STEPQuant is a spatial-temporal post-training quantization framework for Delta-rule recurrent states, achieving over 5x state compression and reducing serving memory by up to 68.7% with minimal accuracy loss.
LeapQuant: Near-Lossless 8-Bit Recurrent State Quantization for Linear Attention LLMs
LeapQuant:突破線性注意力瓶頸,實現近乎無損的 8-bit 遞迴狀態量化
LeapQuant is a training-free quantization method that achieves near-lossless 8-bit recurrent state quantization for linear attention, delivering up to 1.47x end-to-end inference speedup while preserving FP32 accuracy.