A neural network component that learns to match spatial regions in user masks with corresponding appearance features from text or image inputs.