Vision-language models used to evaluate AI agent trajectories are unreliable judges with systematic leniency bias—but open-source alternatives trained on curated data can match commercial performance at a fraction of the cost.
OSReward is a benchmark and dataset for evaluating how well vision-language models can judge whether computer-using AI agents completed tasks correctly. The researchers found that popular VLM judges have a systematic bias toward being too lenient, and released OS-Shepherd, open-source reward models that match expensive commercial judges at 30-60% lower cost.