Visual skills outperform text-based skill representations for embodied agents because they preserve spatial structure; coupling skill learning with policy optimization creates a virtuous cycle of mutual improvement.
ViSkill teaches vision-language model agents to learn and reuse visual skills from successful interactions. Instead of converting spatial information into text (which loses important details), the system stores skills as visual cards that guide both action selection and reward signals, creating a feedback loop where better policies improve the skill library and vice versa.