VCSD achieves significant improvements in vision-language model performance (up to 5% on benchmarks) by using contrastive image conditioning during self-distillation, requiring no external teachers, privileged data, or extra inference costs.
This paper introduces Visual Contrastive Self-Distillation (VCSD), a training method that improves vision-language models by having them learn from themselves without needing an external teacher.