1 citations · 1 across the 4 of their papers we have counts for
7 papers · 1 filter
What Should a Streaming Video Model Remember?
Haonan Ge, Yiwei Wang, Hang Wu +1
Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation bu…
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
Hang Wu, Sherin Mary Mathews, Yujun Cai +2
Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded stream and unpredictable quer…
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
Hang Wu, Yujun Cai, Zehao Li +4
Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification…
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
Bowen Sun, Yujun Cai, Ming-Hsuan Yang +2
Video LLMs suffer from temporal inconsistency: small shifts in frame timing can flip attention and suppress relevant frames. We trace this instability to the common extension of Ro…
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
Hongxing Li, Dingming Li, Zixuan Wang +7
Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We id…
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
Haonan Ge, Yiwei Wang, Kai-Wei Chang +2
Current video understanding models rely on fixed frame sampling strategies, processing predetermined visual inputs regardless of the specific reasoning requirements of each questio…