Sparse autoencoders can turn multimodal model internals into controllable feature interfaces: you can identify what changed during multimodal training, find features causing specific behaviors, and steer or remove them to improve safety or task performance.
This paper introduces MMDiff, a framework that uses sparse autoencoders to decompose multimodal language models into interpretable features.