Aivora

#foil

Foil

1 article

How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention
arXivAI Research

How to Loop MoE: Foil Architecture Flattens Experts and Unties Attention

如何循環混合專家模型?全新 Foil 架構實現專家扁平化與注意力機制解耦

This paper introduces Foil, a novel framework bridging Looped Transformers and sparse MoEs by flattening expert layers and untying attention parameters to boost training efficiency and routing quality.

2 min read