Active visual operations like cropping or zooming that multimodal models perform to gather visual evidence before answering.