Using physical trajectory data as training supervision helps video embedding models better understand motion-centric driving events, improving retrieval accuracy by 5-10% while keeping inference simple and efficient.
This paper tackles retrieving relevant driving video clips from large datasets by improving multimodal embedding models. The key innovation is TraVEL, which fine-tunes video embeddings using trajectory (vehicle motion) as training supervision, helping the model understand motion-centric events like turning or accelerating.