Language phrasing sensitivity in VLAs is a learnable, systematic problem that can be mitigated at inference time using LLM-distilled rephrasing rules, improving robustness to out-of-distribution instructions without model retraining.
Vision-language-action models (VLAs) that control robots are surprisingly fragile to how instructions are phrased—changing one word can drop success rates by 50+ points.