Multimodal agents struggle to couple exploration and visual reasoning in 3D worlds—they can see anomalies or navigate, but struggle to do both together effectively, suggesting a fundamental gap in how these systems integrate action and perception.
WorldAuditBench is a benchmark for testing how AI agents find problems in 3D virtual worlds—like floating objects or walls you can walk through. It evaluates multimodal AI systems (vision-language models and vision-language-action models) on 213 anomaly detection tasks across 13 environments, measuring how well agents can explore systematically and visually identify issues.