Spatially grounded self-distillation with synthetic data can teach multimodal models better visual reasoning that transfers to real-world tasks, without requiring human annotations or external teachers.
This paper improves multimodal AI models by having them learn from a smarter version of themselves that receives spatial hints about where to look in images. Using synthetic scenes with automatic object labels, the approach trains models to understand spatial relationships without human annotation, and surprisingly, these improvements transfer to real-world vision tasks.