Training a model to predict the next frame from previous frames, learning causal temporal structure without action labels.