Vision-language models have a circuit-level sensitivity to input order that can be fixed with simple test-time adaptation, improving both consistency and accuracy without retraining.
Vision-language models perform differently depending on whether you show them the image or question first—a semantically irrelevant change that shouldn't matter. This paper identifies this "modality order" failure and proposes a test-time training method that makes models consistent across both orderings while improving overall performance.