A machine learning approach that processes both audio and visual information together to better understand speech and communication.
Quality of vision, audio, and image understanding (distinct from modality support)