Vision-language models are far better at describing visible spatial relationships than predicting unseen outcomes after changes—a gap that matters for real-world applications like robotics and planning.
SpaceCast-Bench is a new benchmark that tests whether vision-language models can predict how scenes change after interventions, not just describe what they see. With 3,862 questions from real scenes, it reveals that even top models reach only 58% accuracy versus 87% for humans, and that models struggle without explicit 3D information or bridge views connecting different observations.