Aivora
arXivLLMAdvanced

Scaling Laws for Looped Mixture of Experts: Combining Recurrence and Sparsity

融合循環與稀疏架構:Looped MoE 的全新縮放定律

2 min read
Scaling Laws for Looped Mixture of Experts: Combining Recurrence and Sparsity
The 30-second version

This paper presents the Loop Scaling Laws, the first framework to jointly model recurrence and sparsity alongside model size and data. While recurrence increases computational depth without adding parameters, MoE expands model capacity without increasing active compute. The study shows that sparsity yields ~3x active-parameter efficiency, and recurrence delivers ~2x total-parameter efficiency on reasoning tasks. At a trillion-token scale, a Looped MoE matches the performance of a non-looped MoE twice its size, while enabling dynamic test-time scaling.

Key points

01

First Joint Scaling Law

Successfully unifies recurrence and MoE sparsity into a single mathematical framework, predicting held-out loss more accurately than prior methods.

02

Complementary Efficiency Gains

Sparsity delivers ~3x active-parameter efficiency, while recurrence yields ~2x total-parameter efficiency on complex reasoning tasks.

03

Trillion-Token Validation

At a trillion-token scale, a Looped MoE matches a ~2x larger non-looped model under matched compute, enabling dynamic test-time scaling.

How it works

Model Comparison: Traditional vs. Looped MoE
Dense (稠密模型)Standard MoE (標準 MoE)Looped MoE (迴圈 MoE)
Active Param Efficiency基準 (1x)高 (~3x 效率)高 (~3x 效率)
Total Param Efficiency基準 (1x)基準 (1x)極高 (~2x 推理效率)
Test-Time Scaling不支援不支援支援 (調整遞迴次數)

Why it matters

This research breaks the bottleneck of traditional model scaling by showing that recurrence and MoE are highly synergistic. It provides a principled design foundation under tight memory and compute constraints. Developers can now build models with smaller memory footprints that still leverage the capacity of MoE. The ability to perform test-time scaling via variable recurrence steps opens up new avenues for efficient inference and adaptable deployment.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager

How to use it

  1. 1Resource-Constrained Deployment: Running high-performance LLMs on edge devices or consumer hardware with strict memory and compute limits.
  2. 2Test-Time Dynamic Scaling: Dynamically adjusting recurrence steps during inference based on task complexity to optimize operational compute costs.

Limitations & caveats

  • Recurrent architectures may increase sequential inference latency due to the serial nature of deep loops, limiting parallelization.
  • The scaling law parameters may require re-calibration when applied to highly heterogeneous datasets or novel MoE routing mechanisms.

Related

SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability
arXivLLM

SCAPO: Optimizing Token-Level Credit in RLVR via Semifactual Stability

SCAPO:藉由半事實穩定性最佳化 RLVR 的 Token 級信用分配

SCAPO is a novel variant of GRPO that incorporates semifactual stability into token-level credit assignment, significantly improving LLM reasoning accuracy and out-of-distribution generalization.

2 min read
STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States
arXivLLM

STEPQuant: Spatial-Temporal Quantization for Linear Attention Recurrent States

STEPQuant:空間與時間雙重優化,突破線性注意力循環狀態量化瓶頸

STEPQuant is a spatial-temporal post-training quantization framework for Delta-rule recurrent states, achieving over 5x state compression and reducing serving memory by up to 68.7% with minimal accuracy loss.

2 min read