Current VLA models fail on complex spatial reasoning and multi-step planning—RoboSPA provides a diagnostic benchmark showing where embodied AI systems need improvement beyond simple, short-horizon tasks.
RoboSPA is a large-scale robotic manipulation dataset and benchmark designed to test how well Vision-Language-Action (VLA) models handle complex spatial reasoning and long-horizon planning tasks. It includes 527K trajectories across 280 task variants with increasing difficulty, revealing that current VLA models struggle with spatial relations, precise execution, and memory-intensive planning.