Planning in a world model's latent noise space instead of raw actions significantly improves closed-loop control performance, and decoder-free latent representations enable better state alignment than reconstruction-based approaches.
LeWAM is a world action model that predicts both actions and future states using a latent representation from JEPA (a self-supervised learning approach). Unlike models that reconstruct images, LeWAM works directly in a learned latent space and can predict forward, backward, and inverse dynamics.