Training MLLMs to internally reconstruct 3D scene geometry (even in compact form) improves spatial reasoning and 3D understanding, suggesting that learning to imagine scenes is more effective than explicit geometric supervision alone.
This paper teaches multimodal AI models to reason about 3D scenes by first imagining a compact 3D representation before answering questions. Instead of relying on detailed geometric details, the model learns to assemble a coarse 3D layout from multiple viewpoints—similar to how humans understand 3D space—then uses this mental model to answer spatial reasoning questions more accurately.