For self-training, sampling diverse problem-solving strategies matters more than correctness or teacher model size—a small model trained on varied approaches beats distillation from a 235B teacher.
This paper shows that self-training works better when you sample diverse problem-solving approaches rather than just correct answers. The authors introduce GROOT (a tree-based sampling method) and Verbalized Sampling to generate strategically different solutions, and find that models trained on diverse but incorrect traces outperform those trained on correct answers from much larger models.