By attributing teacher corrections to visual evidence through counterfactual intervention, VAD extracts purer visual supervision signals for distillation, outperforming naive privileged-view supervision that mixes visual and linguistic signals.
This paper addresses a key problem in multimodal distillation: when a teacher corrects a student's mistakes, it's unclear how much of that correction comes from visual evidence versus linguistic priors. VAD solves this by using counterfactual reasoning—removing visual evidence and measuring how the teacher's predictions change—to isolate visually-grounded corrections.