When combining multiple RL-trained teachers into one student, the averaging method and optimizer choice matter more than raw gradient differences—response length weighting and precision loss can swing task performance by 2-5 percentage points.
This paper investigates how multiple teacher models transfer knowledge to a student model during on-policy distillation. The researchers found that loss averaging implicitly weights responses, Adam's optimizer smooths gradient differences, and low-precision arithmetic (BF16) masks small weight updates—with these factors significantly affecting which tasks the student learns best.