You can improve robot manipulation models through human feedback without robot execution by using a handheld interface to detect when the policy is uncertain and identifying which parts of demonstrations are most important for learning.
This paper presents HIL-UMI, a method for improving vision-language-action robot models without needing a physical robot during training. Instead of repeatedly running the robot to collect new data, humans demonstrate tasks using a handheld interface while the system queries the current policy and intelligently decides when to collect new examples based on policy uncertainty and task progress.