Predicting learned motion representations instead of raw joint commands lets humanoid policies leverage massive human motion datasets for zero-shot real-world control, solving the dual problems of high-dimensional action spaces and scarce robot demonstrations.
VioLA is a humanoid robot control policy that learns from human motion data by predicting body and hand motion patterns instead of direct joint commands. By training on 140 million frames (93% human data), it achieves zero-shot task execution on real robots without task-specific fine-tuning, reaching 100% success on locomotion where prior methods scored 0-17%.