On-policy distillation wastes data but struggles with learning speed: a single query exposes most supervision needed, but the student takes hundreds of training steps to absorb it, suggesting future improvements should focus on faster alignment rather than more data.
This paper investigates how much training data on-policy distillation (OPD) actually needs by training on just a single query. Surprisingly, one query recovers most of the performance gains of full-dataset training, reaching 71.5% of the state coverage that full data achieves.