Privileged information in self-distillation matters less than the cross-mode transfer between direct-response and thinking-enabled inference; most gains come from distillation itself, not from what the reference reveals.
This paper investigates what privileged information (like worked solutions or reasoning traces) actually adds to on-policy self-distillation in language models.