Video models need different scaling strategies than language models—splitting experts by semantic role and allowing imbalanced routing produces better video generation than forcing uniform expert usage.
This paper proposes SplitMoE, a new way to scale video generation models using a split expert architecture that avoids forcing uniform expert usage.