Natural language is becoming a primary feedback mechanism for training and guiding AI agents—moving beyond traditional numerical rewards to leverage language's ability to convey intent, preferences, and reasoning in ways both humans and LLMs understand.
This paper introduces Verbal Reinforcement Learning (VRL), a framework where natural language serves as feedback to improve language agents. It organizes the field into three categories: language defining tasks and rewards, language guiding reasoning at test time, and language shaping model parameters during training.