Anchoring model reasoning to specific locations and moments in video, connecting visual content to semantic meaning.