The process of identifying and converting text that appears in images or documents into machine-readable text format.
Quality of vision, audio, and image understanding (distinct from modality support)