A machine learning model trained to understand and process images as input.
Quality of vision, audio, and image understanding (distinct from modality support)