This benchmark reveals that gameplay requires coordinating visual understanding, instruction following, and long-horizon planning—and shows clear performance gaps between model families, providing a standardized way to measure progress on embodied AI tasks.
GameHorizon is a comprehensive dataset and benchmark for evaluating AI models on video game tasks. It includes 5,000 hours of gameplay from 21 AAA games with aligned videos, actions, and instructions at multiple time scales, plus offline and online evaluation tracks to measure how well models understand and execute gameplay across different planning horizons.