By interleaving visual objects directly into text during pretraining, you can teach multimodal models object-level grounding 12x more efficiently than traditional image-text pair training.
This paper introduces MultiModal Code-Switching (MMCS), a new way to train vision-language models by replacing words in text with their corresponding visual objects. Instead of just pairing whole images with descriptions, MMCS explicitly shows the model which objects match which words, making training much more efficient—achieving the same performance with 12x less data.