DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models — ThinkLLM