The degree to which a model produces similar correct or incorrect answers across different input formats like text and images.