long video processing 1multimodal retrieval 1sparse token selection 1vision-language models 1visual retrieval 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…
cs.CV2026
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
Yuqun Wu, Chih-hao Lin, Henry Che +4
We investigate the problem of identifying objects that have been added, removed, or moved between a pair of captures (images or videos) of the same scene at different times. Accura…
cs.CV2025
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
Yao Xiao, Qiqian Fu, Heyi Tao +3
Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like…