You can efficiently learn video representations by repurposing pretrained image models with clever masking strategies, avoiding expensive 3D architectures and reconstruction overhead.
VideoMSN uses standard image Vision Transformers to learn video representations without 3D models or reconstruction. It treats videos as grids of frames and masks either spatial patches or temporal frames, then aligns the two views using a Siamese loss. This achieves top results on video benchmarks while needing 32-160x fewer training epochs than prior methods.