You can now use different models to generate training data and provide distillation supervision without performance loss, by filtering out style differences rather than requiring teacher consistency.
This paper addresses a practical problem in training large reasoning models: when the teacher model used for distillation differs from the one that generated the training data, performance often suffers.