You can improve LLMs through self-distillation using only the model's own outputs and internal consistency, without needing ground-truth labels or external feedback—making it truly self-supervised post-training.
This paper introduces Unsupervised On-Policy Self-Distillation (U-OPSD), a method that improves language models without requiring external supervision like ground-truth answers or feedback from larger models.