Agents can learn to act better by learning to explain their actions—training on self-generated retrospections alone improves future performance without RL, suggesting explanation is a useful learning signal for behavior improvement.
This paper shows that language model agents can improve their performance by training on self-generated explanations of their own experiences, without needing reinforcement learning or external rewards. The method, called Retrospection-Only Fine-Tuning (ROFT), has an agent attempt tasks, generate explanations of what happened, and then fine-tune on predicting those explanations.