Skills improve LLM agents less through new capabilities and more through avoiding regressions. Better agent reliability comes from improving how skills ground inputs and verify outputs, not from adding more procedural steps.
When you add skills to an AI agent, it doesn't just help—it can also hurt performance on tasks it previously solved. This paper measures both gains and losses across 6,000 test runs, finding that top-performing skills succeed mainly by causing fewer regressions, not by enabling more wins.