A single trainable embedding can dramatically improve how vision-language models handle long visual context by acting as a retrieval target—selecting sparse, relevant tokens instead of processing everything, making it practical for real-world deployment.
ReToken is a lightweight technique that adds a single learnable token to vision-language models to intelligently select relevant visual information from long images and videos.