Aivora
arXivAI ResearchAdvanced

How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention

如何循環混合專家模型?全新 Foil 架構實現專家扁平化與注意力機制解耦

2 min read
How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention
The 30-second version

Looped Transformers reuse layers to maximize parameter utilization, while MoE limits computational cost per token. Combining them (Looped MoE) is highly promising but challenging. The authors propose Foil, which optimizes looped MoE under fixed parameters and compute. Foil (1) "flattens" the experts by halving expert layers, doubling experts per layer, and doubling the passes; and (2) "unties" the attention by assigning separate attention parameters to each pass while sharing experts/routers. At 100B tokens, Foil outperforms unflattened baselines, achieving lower pretraining loss and more balanced routing.

Key points

01

Bridging Looped and MoE

Combines the benefits of Looped Transformers (layer reuse) and MoE (sparse activation) to maximize parameter efficiency without increasing compute.

02

Flattening the Experts

Halves expert layers, doubles experts per layer, and doubles the passes, allowing each routing decision to choose from a larger pool of experts.

03

Untying Attention

Equips each pass with its own unique attention parameters while keeping experts and routers shared, yielding more balanced and confident routing.

04

Lower Pretraining Loss

At 100B tokens, the most flattened Foil reduces pretraining loss by 0.012 nat compared to the baseline, maintaining or improving downstream accuracy.

How it works

Comparison between Foil Architecture and Baseline Looped MoE
Baseline (Unflattened Looped MoE)Foil (Flattened & Untied MoE)
Expert Layers較多層(正常層數)層數減半(扁平化)
Experts per Layer較少專家專家數量加倍
Number of Passes基礎循環次數循環次數加倍
Attention Parameters在循環間完全共享 (Shared)每次循環配備獨立參數 (Untied)
Experts & Routers共享保持共享

Why it matters

Scaling up LLMs traditionally demands prohibitive memory and compute. Looped MoE offers a path toward highly parameter-efficient yet high-capacity models. By proving that structural modifications—such as flattening experts and untying attention—can significantly boost performance under fixed compute, Foil provides a practical, tested blueprint for designing next-generation lightweight, highly-efficient AI models.

Who it affects

  • AI Researcher
  • AI Developer

How to use it

  1. 1Efficient LLM deployment in memory-constrained environments
  2. 2Maximizing pretraining and inference efficiency under strict compute budgets

Limitations & caveats

  • The doubled number of passes may increase sequential inference latency.
  • While attention is untied, the highly shared expert network structure requires further validation at extreme model scales.

Related

FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
arXivAI Research

FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets

FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術

FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.

2 min read