collaborators

11 papers

cs.CV2026

What Should a Streaming Video Model Remember?

Haonan Ge, Yiwei Wang, Hang Wu +1

Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation bu…

cs.CV2026

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

Hang Wu, Sherin Mary Mathews, Yujun Cai +2

Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded stream and unpredictable quer…

cs.CV2026

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning

Hang Wu, Yujun Cai, Zehao Li +4

Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification…

cs.CV2026

PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs

Bowen Sun, Yujun Cai, Ming-Hsuan Yang +2

Video LLMs suffer from temporal inconsistency: small shifts in frame timing can flip attention and suppress relevant frames. We trace this instability to the common extension of Ro…

cs.RO2026

StreamVLA: Breaking the Reason-Act Cycle via Completion-State Gating

Tongqing Chen, Hang Wu, Jiasen Wang +2

Long-horizon robotic manipulation requires bridging the gap between high-level planning (System 2) and low-level control (System 1). Current Vision-Language-Action (VLA) models oft…

cs.CV2025

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models

Hongxing Li, Dingming Li, Zixuan Wang +7

Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We id…