How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention
如何循環混合專家模型?全新 Foil 架構實現專家扁平化與注意力機制解耦
Looped Transformers reuse layers to maximize parameter utilization, while MoE limits computational cost per token. Combining them (Looped MoE) is highly promising but challenging. The authors propose Foil, which optimizes looped MoE under fixed parameters and compute. Foil (1) "flattens" the experts by halving expert layers, doubling experts per layer, and doubling the passes; and (2) "unties" the attention by assigning separate attention parameters to each pass while sharing experts/routers. At 100B tokens, Foil outperforms unflattened baselines, achieving lower pretraining loss and more balanced routing.
Key points
Bridging Looped and MoE
Combines the benefits of Looped Transformers (layer reuse) and MoE (sparse activation) to maximize parameter efficiency without increasing compute.
Flattening the Experts
Halves expert layers, doubles experts per layer, and doubles the passes, allowing each routing decision to choose from a larger pool of experts.
Untying Attention
Equips each pass with its own unique attention parameters while keeping experts and routers shared, yielding more balanced and confident routing.
Lower Pretraining Loss
At 100B tokens, the most flattened Foil reduces pretraining loss by 0.012 nat compared to the baseline, maintaining or improving downstream accuracy.
How it works
| Baseline (Unflattened Looped MoE) | Foil (Flattened & Untied MoE) | |
|---|---|---|
| Expert Layers | 較多層(正常層數) | 層數減半(扁平化) |
| Experts per Layer | 較少專家 | 專家數量加倍 |
| Number of Passes | 基礎循環次數 | 循環次數加倍 |
| Attention Parameters | 在循環間完全共享 (Shared) | 每次循環配備獨立參數 (Untied) |
| Experts & Routers | 共享 | 保持共享 |
Why it matters
Scaling up LLMs traditionally demands prohibitive memory and compute. Looped MoE offers a path toward highly parameter-efficient yet high-capacity models. By proving that structural modifications—such as flattening experts and untying attention—can significantly boost performance under fixed compute, Foil provides a practical, tested blueprint for designing next-generation lightweight, highly-efficient AI models.
Who it affects
- AI Researcher
- AI Developer
How to use it
- 1Efficient LLM deployment in memory-constrained environments
- 2Maximizing pretraining and inference efficiency under strict compute budgets
Limitations & caveats
- The doubled number of passes may increase sequential inference latency.
- While attention is untied, the highly shared expert network structure requires further validation at extreme model scales.
Related
FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術
FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.
Self-Correcting Multimodal Models: UMM-Reflection Enables Native Image Generation Repair via Interleaved RL
讓多模態模型自我修正!UMM-Reflection 透過交錯強化學習實現原生圖像生成反思
UMM-Reflection introduces interleaved reinforcement learning to enable a single unified multimodal model to self-diagnose and repair its own generated images without external verifiers.
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
Researchers found that training reasoning models to predict their own confidence at intermediate steps naturally reduces generated tokens by up to 25% at matched accuracy, without explicitly optimizing for length or stopping.