By distilling VLM-identified task-salient information into a learned latent token during training, robots can efficiently handle long-horizon tasks at deployment time without expensive in-the-loop reasoning.
This paper introduces workspace tokens, a lightweight memory system for robotic manipulation that learns which task-relevant information to remember during training using a vision-language model, then uses this compressed memory at deployment without needing expensive VLM queries.