A training method where the image encoder is frozen and kept unchanged while only the text processing components are trained.
Quality of vision, audio, and image understanding (distinct from modality support)