Token-level stability under prompt variations is a useful training signal for improving reasoning in LLMs—you can boost performance by penalizing tokens that change meaning when irrelevant prompt details change.
This paper identifies that language models trained with reinforcement learning are sensitive to irrelevant prompt changes, even when the problem stays the same. The authors propose SCAPO, an improved training method that assigns credit to individual tokens based on how stable they are under these prompt variations.