Video models fail not because they can't see events, but because they can't reliably track and count them over time—adding more frames helps slightly but doesn't fix the core temporal reasoning problem.
Video language models struggle with counting events in videos, especially when events happen frequently or repeatedly.