Adaptive weighting of token-level supervision based on divergence sequences improves reasoning model training—the timing and context of prediction errors matter more than their magnitude alone.
This paper improves on-policy self-distillation for reasoning models by making supervision weights adaptive to how prediction errors evolve during generation. Instead of treating all mismatches between teacher and student equally, DASH uses gates to weight supervision based on local divergence patterns, improving math reasoning across model sizes without extra computation.