When building multimodal models, image tokenizer choice matters not just for image quality but for how well it helps the model learn text and vision together—and the best tokenizer for one task may not be best for another.
This paper studies how image tokenizers work in multimodal AI models by building a controlled training setup that tracks how well different tokenizers perform across text, image generation, and image understanding tasks.