Using a VLM to detect actual visual transitions in videos produces better event localization and captions than pre-written transition descriptions inserted at fixed positions.
This paper tackles dense video captioning—describing multiple events in long videos—by using a vision-language model to intelligently detect transition moments between events rather than blindly inserting captions everywhere.