Object-centric predictive models trained on multi-view video can enable underwater robots to plan manipulation tasks by imagining future states, without requiring expensive contact sensors or reconstruction.
This paper presents Underwater C³-JEPA, a world model that predicts how objects move during underwater robot salvage operations. Using multiple camera views and robot control signals, it learns to predict future object positions in a compact latent space without needing contact sensors. The model transfers well to downstream tasks and works on real underwater video.