Distilling ranking knowledge from a reranker into embedding models significantly improves compositional reasoning—the ability to distinguish scenes by attribute-object relationships—achieving 82.7% on compositional benchmarks while maintaining standard retrieval performance.
This paper improves how multimodal AI models understand complex visual scenes by teaching embedding models to better distinguish between scenes with the same objects but different arrangements.