Combining generated video with generated audio allows robots to infer not just motion but also the forces needed for contact-heavy tasks—something video alone cannot provide.
This paper shows how robots can learn contact-rich manipulation tasks by generating both video and audio together. The system uses the loudness of generated contact sounds to create force profiles that guide the robot's movements, enabling tasks like pushing and grasping that require precise force control. The approach works without task-specific training data.