To build better looped MoE models, flatten the expert hierarchy (more experts per layer, more passes) and give each pass independent attention—this lets tokens access more experts while maintaining computational efficiency.
This paper improves looped mixture-of-experts (MoE) models by flattening the expert structure and untying attention parameters. The key insight is that by doubling experts per layer and doubling passes through the network while keeping compute fixed, models can route tokens through more diverse experts, improving performance.