arXivAI Research
How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention
如何循環混合專家模型?全新 Foil 架構實現專家扁平化與注意力機制解耦
This paper introduces Foil, a novel framework bridging Looped Transformers and sparse MoEs by flattening expert layers and untying attention parameters to boost training efficiency and routing quality.
2 min read