VLA robotic models share cross-task vulnerabilities that can be exploited with a single adversarial texture, revealing a critical safety gap in multitask embodied AI systems that current defenses don't address.
This paper demonstrates how a single adversarial texture on a 3D object can fool vision-language-action (VLA) robotic models across multiple tasks simultaneously. Rather than crafting separate attacks for each task, the researchers optimize one texture that works universally by backpropagating gradients through a differentiable renderer, reducing task success rates from 90% to 48% in experiments.