End-to-end vision-language models can control complex humanoid robots for real-world navigation by learning whole-body coordination in simulation, eliminating the need for modular planning pipelines and real-world data collection.
TANGO is a vision-language model that enables humanoid robots to navigate cluttered indoor spaces by predicting full-body joint movements directly from natural language instructions and camera images. Unlike traditional 2D path planning, it coordinates arm placement, torso adjustment, and walking patterns to move through complex 3D environments.