Vision-language reward models for robotics are fragile to paraphrasing—rewording the same goal can flip success/failure judgments on identical robot behavior, a critical flaw for reliable robotic learning systems.
Vision-language models are being used to score robot behavior, but they fail a basic requirement: giving the same score when instructions are paraphrased. This paper introduces ROBORMBENCH, a benchmark of 2,390 real robot trajectories with 21,673 paraphrases, showing that current VLMs flip between calling identical robot actions successful or failed depending on how you word the goal.