Humanoid robots struggle with tool selection and coordinating manipulation with movement—even state-of-the-art models like GR00T show reduced accuracy on unseen tools and can execute tasks despite receiving unrelated instructions.
This paper introduces HumanoidToolBench, a benchmark for evaluating humanoid robots on tool use tasks that require selecting appropriate tools and coordinating manipulation with locomotion. The benchmark includes 3,100 demonstrations and tests seven policies, revealing significant gaps between tool selection and successful task completion.