Models struggle to switch between different reasoning skills in long tasks, and explicitly training them to recognize and predict which skill to use at each step dramatically improves their ability to handle complex multi-step reasoning.
This paper introduces Skill Entropy, a metric for measuring how difficult it is for language models to switch between different reasoning skills (like math, planning, coding) in long multi-step tasks. The authors create Skill²-Bench, a benchmark with 558 skills across 9 domains, and show that current models struggle more on tasks requiring frequent skill switches.