13 citations · 18 across the 23 of their papers we have counts for
5 papers · 2 filters
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infe…
Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents
Dong-Hee Kim, Reuben Tan, Donghyun Kim
Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied these tools in visual search ta…
InstrAct: Towards Action-Centric Understanding in Instructional Videos
Zhuoyi Yang, Jiapeng Yu, Reuben Tan +2
Understanding instructional videos requires recognizing fine-grained actions and modeling their temporal relations, which remains challenging for current Video Foundation Models (V…
Learning Sparse Visual Representations via Spatial-Semantic Factorization
Theodore Zhengde Zhao, Sid Kiblawi, Jianwei Yang +6
Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens th…
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
Zane Durante, Silky Singh, Arpandeep Khatua +6
Training video-language models is often prohibitively expensive due to the high cost of processing long frame sequences and the limited availability of annotated long videos. We pr…