11 papers
What Should a Streaming Video Model Remember?
Haonan Ge, Yiwei Wang, Hang Wu +1
Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation bu…
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
Hang Wu, Sherin Mary Mathews, Yujun Cai +2
Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded stream and unpredictable quer…
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
Hang Wu, Yujun Cai, Zehao Li +4
Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification…
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
Bowen Sun, Yujun Cai, Ming-Hsuan Yang +2
Video LLMs suffer from temporal inconsistency: small shifts in frame timing can flip attention and suppress relevant frames. We trace this instability to the common extension of Ro…
StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating
Tongqing Chen, Hang Wu, Jiasen Wang +2
Long-horizon robotic manipulation requires bridging the gap between high-level planning (System 2) and low-level control (System 1). Current Vision-Language-Action (VLA) models oft…
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
Hongxing Li, Dingming Li, Zixuan Wang +7
Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We id…