Decoupling exploration (with novelty bonuses) from optimization (standard training) via distillation lets language models discover diverse correct reasoning strategies without quality degradation, outperforming direct RLVR approaches.
This paper proposes Exploration-Distillation (ExpDis), a method that separates exploration from optimization in reinforcement learning with verifiable rewards. Explorer policies use novelty bonuses to discover new reasoning strategies, their best trajectories are filtered and distilled into a student policy trained without novelty incentives.