Visual distractors cause robot policies to select wrong objects not because they lose manipulation skills, but because they struggle to ground attention on the correct target—a problem that can be fixed by explicitly training better target selection.
This paper identifies why robot learning policies fail when similar-looking objects are present, even though they work well normally.