Looped MoE models combine two orthogonal efficiency axes: recurrence increases computational depth while sparsity expands capacity, and their scaling laws enable principled design of efficient models that match much larger dense models at the same compute budget.
This paper develops scaling laws that jointly model looped transformers (which use recurrence for computational depth) and Mixture-of-Experts (which use sparsity for capacity).