collaborators

8 papers

cs.CV2026

TimeThink: Reasoning with Time for Video LLMs

Handong Li, Longteng Guo, Zikang Liu +8

Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promisi…

cs.CV2026

Conditional Multi-Event Temporal Grounding in Long-Form Video

Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez +12

Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositi…

cs.CV2026

PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding

Nan Wang, Zhiwei Jin, Chen Chen +1

Document understanding and GUI interaction are among the highest-value applications of Vision-Language Models (VLMs), yet they impose exceptionally heavy computational burden: fine…

cs.CV2026

STAR: Mitigating Cascading Errors in Spatial Reasoning via Turn-point Alignment and Segment-level DPO

Pukun Zhao, Longxiang Wang, Chen Chen +4

Structured spatial navigation is a core benchmark for Large Language Models (LLMs) spatial reasoning. Existing paradigms like Visualization-of-Thought (VoT) are prone to cascading…

cs.CV2025

From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding

Guangyu Sun, Archit Singhal, Burak Uzkent +3

Video Large Language Models (VLMs) have achieved strong performance on various vision-language tasks, yet their practical use is limited by the massive number of visual tokens prod…

cs.CL2025

Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)

Jing Bi, Susan Liang, Xiaofei Zhou +16

Reasoning is central to human intelligence, enabling structured problem-solving across diverse tasks. Recent advances in large language models (LLMs) have greatly enhanced their re…