Failed expert trajectories are learning gold: teaching models to reflect on why attempts failed is often easier than solving hard problems from scratch, and this reflective skill transfers to direct problem-solving.
ReflectRL shows that when expert models fail on hard problems, their failed trajectories aren't wasted—they're valuable for learning. The method trains models to reflect on these flawed attempts, then transfers that reflective reasoning back to solving problems directly. This lightweight approach works across multiple benchmarks and training methods with minimal computational cost.