When multimodal models favor one input type (e.g., text) over others (e.g., vision), bypassing available tools.