activity
20192026
most citedLook at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

13 citations · 18 across the 23 of their papers we have counts for

collaborators
Showing 2026 · cs.CVShow all

5 papers · 2 filters

cs.CV2026

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Yao Xiao, Reuben Tan, Zhen Zhu +3

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infe…

cs.CV2026

Diversity Over Frequency: Rethinking Tool Use in Visual Chain-of-Thought Agents

Dong-Hee Kim, Reuben Tan, Donghyun Kim

Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied these tools in visual search ta…

cs.CV2026

InstrAct: Towards Action-Centric Understanding in Instructional Videos

Zhuoyi Yang, Jiapeng Yu, Reuben Tan +2

Understanding instructional videos requires recognizing fine-grained actions and modeling their temporal relations, which remains challenging for current Video Foundation Models (V…

cs.CV2026

Learning Sparse Visual Representations via Spatial-Semantic Factorization

Theodore Zhengde Zhao, Sid Kiblawi, Jianwei Yang +6

Self-supervised learning (SSL) faces a fundamental conflict between semantic understanding and image reconstruction. High-level semantic SSL (e.g., DINO) relies on global tokens th…

cs.CV2026

VideoWeave: A Data-Centric Approach for Efficient Video Understanding

Zane Durante, Silky Singh, Arpandeep Khatua +6

Training video-language models is often prohibitively expensive due to the high cost of processing long frame sequences and the limited availability of annotated long videos. We pr…