By organizing self-improvement around dynamically managed skills, LLMs can achieve both reliable feedback and open-ended task diversity—enabling more robust self-evolution than existing methods.
This paper introduces Skill Self-Play, a framework where language models improve themselves through co-evolving components: a task proposer, a solver, and a skill controller.