Today's multimodal AI models lack active visual observation—they don't strategically re-examine images to gather information, a core capability humans use for vision tasks. This gap persists even when models write their own vision code.
ActiveVision is a benchmark that tests whether multimodal AI models can actively observe images by repeatedly examining different parts, similar to how humans use eye movements. Current frontier models like GPT-4o and Claude fail dramatically (3-10% accuracy) compared to humans (96%), revealing that these models treat images as static snapshots rather than actively exploring them to solve tasks.