collaborators

6 papers

cs.CV2026

video-SALMONN-R: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

Yixuan Li, Guangzhi Sun, Yudong Yang +1

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to…

cs.CV2026

Holo-World: Unified Camera, Object and Weather Control for Video World Model

Xiangchen Yin, Wenzhang Sun, Jiahui Yuan +6

Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls r…

cs.CV2026

A DVDrive Approach for doScenes Instructed Driving Challenge

Zijian Fu, Xiangyang Chu, Mengshi Qi +3

Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and…

cs.CV2026

GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking

Zixu Cheng, Da Li, Jian Hu +4

Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language…

cs.CV2025

V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

Zixu Cheng, Jian Hu, Ziquan Liu +3

Human processes video reasoning in a sequential spatio-temporal reasoning logic, we first identify the relevant frames ("when") and then analyse the spatial relationships ("where")…

cs.CV2025

CoS: Chain-of-Shot Prompting for Long Video Understanding

Jian Hu, Zixu Cheng, Chenyang Si +2

Multi-modal Large Language Models (MLLMs) struggle with long videos due to the need for excessive visual tokens. These tokens exceed massively the context length of MLLMs, resultin…