VLMs have a critical weakness: they often trust learned world knowledge over what's actually shown in images, and SABRE provides a reusable framework to systematically identify and measure such failures.
SABRE is an automated pipeline that creates stress tests for vision-language models by converting task designs into images and question-answer pairs. It tests whether VLMs rely on visual evidence or learned assumptions about the world, revealing that current models struggle significantly (17.8-31.3% accuracy) when images contradict expectations.