Visual tool-use in multimodal models often creates an illusion of improvement: while aggregate accuracy may increase, the returned visual evidence frequently has no causal effect on the final answer, making these expensive operations ineffective for most queries.
This paper investigates whether multimodal AI models actually benefit from visual tool-use operations like cropping and zooming. Using causal analysis, the authors discover that these operations often don't meaningfully improve answers despite higher computational cost—models either ignore the visual evidence they retrieve or use it incoherently.