Using LLM-generated semantics as a shared anchor point for aligning incomplete multimodal data is more robust than trying to reconstruct missing modalities or design complex fusion mechanisms.
SemMSA tackles multimodal sentiment analysis when some data is missing by using large language models to create rich semantic representations that ground all modalities together. Instead of reconstructing missing features, it aligns visual, acoustic, and text representations through spectral methods, achieving better results on standard benchmarks.