When a model's ability to solve a problem varies depending on whether it receives text, images, or both.