The process of aligning text phrases from a caption with corresponding pixel-level segmentation masks.