By training on cross-embodiment video data with a unified physics understanding, CLAP creates robot world models that match or beat single-embodiment models and generalize to new robots without retraining.
CLAP is a video world model that learns physics from diverse videos of humans and robots by treating physical laws as universal. It solves the challenge of different action representations across embodiments using end-effector poses, language, and learned latent actions, then applies a curriculum approach to train models that work zero-shot on real robot tasks.