VLMs can be made 3D-aware by injecting learned geometric representations from video, enabling strong performance on spatial reasoning tasks while remaining RGB-only—no special 3D sensors or data needed.
This paper enhances vision-language models to better understand 3D spatial information by adding two types of geometric representations learned from RGB videos: implicit geometry tokens that capture high-level 3D structure, and explicit geometry tokens that encode detailed geometric details.