VLMs can be more useful as selective training guides than as direct policies—by querying them only when needed and validating their advice through environment rewards, you can build cheaper, better-performing autonomous agents that don't depend on the VLM at deployment.
This paper presents SAGE, a method for training lightweight autonomous policies by selectively querying an expensive Vision-Language Model (VLM) teacher during training.