Combining dense captioning with pixel-level grounding requires jointly optimizing text generation and mask selection—PANORAMA shows this can be done effectively by conditioning a segmenter on phrase representations and learning which masks correspond to each phrase.
This paper introduces PANORAMA, a vision-language model that generates detailed image captions while simultaneously grounding each phrase with pixel-level segmentation masks.