Separating geometry into common and residual components while routing visual learning through structured latents prevents representation collapse and improves 3D spatial reasoning from images.
This paper addresses 3D spatial reasoning from 2D images by introducing GeoLatent, which uses decomposed spatial representations (position, direction, geometry) with geometric supervision and routed optimization. The method prevents geometry collapse and ensures latents are actively used during learning, achieving state-of-the-art results on spatial reasoning benchmarks.