Visual transition reasoning—understanding how scenes change—is a foundational skill that multimodal LLMs can learn from diverse video examples and reuse across many downstream tasks, offering a systematic way to improve spatial and temporal reasoning.
WOVEN is a training dataset and benchmark for teaching multimodal LLMs to reason about how visual scenes change over time. The authors create 36,000+ examples showing different actions in various scenes, train multiple models on this data, and find that this 'visual transition reasoning' skill transfers to 22 other benchmarks—improving performance by up to 27 percentage points.