Streaming VLMs fail at perceiving fast events in videos—the best model scores only 50.7% on FastBench—because they can't adaptively determine when to sample frames densely, revealing a fundamental gap between sparse uniform sampling and what's actually needed for dynamic understanding.
FastBench is a benchmark for evaluating how well streaming video language models can understand fast-moving events in real-world videos. Current models struggle with this task because they must balance limited context budgets across temporal history, spatial resolution, and frame sampling rates.