Multimodal models can learn better by having strong views teach weaker views on the same problem—using the model's own best performance as supervision rather than external labels.
Vision-language models struggle with visual reasoning even when problems have equivalent text and diagram versions. This paper shows different views (text, image, combined) expose different reasoning paths and failure modes.