Music tokenization design is more important than model size for text-to-music generation—a small model with the right token representation outperforms massive models with poor representations, challenging the field's scaling assumptions.
This paper investigates how music tokenization—the way music is converted into discrete symbols for language models—affects text-to-music generation quality. By fixing model size, data, and training approach while swapping only the tokenization scheme, researchers found that representation choice matters far more than model scale.