Don't trust standard alignment scores as proof that multimodal models understand images—they can be fooled by the model's internal structure. Use task-based validation and the PA gap metric instead to verify genuine cross-modal integration.
This paper reveals that high visual-text similarity scores in multimodal AI models don't actually mean the model is properly integrating images and text. Using experiments across 13 models, the researchers show that replacing images with random noise barely changes similarity scores, even though task performance drops sharply.