6 papers
video-SALMONN-R: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
Yixuan Li, Guangzhi Sun, Yudong Yang +1
Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to…
Holo-World: Unified Camera, Object and Weather Control for Video World Model
Xiangchen Yin, Wenzhang Sun, Jiahui Yuan +6
Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state to change. Yet these controls r…
A DVDrive Approach for doScenes Instructed Driving Challenge
Zijian Fu, Xiangyang Chu, Mengshi Qi +3
Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only from visual scene context and…
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Zixu Cheng, Da Li, Jian Hu +4
Video reasoning requires a fine-grained understanding of the temporal dependencies and event-level relations between objects and events in videos. Current Multimodal Large Language…
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
Zixu Cheng, Jian Hu, Ziquan Liu +3
Human processes video reasoning in a sequential spatio-temporal reasoning logic, we first identify the relevant frames ("when") and then analyse the spatial relationships ("where")…
CoS: Chain-of-Shot Prompting for Long Video Understanding
Jian Hu, Zixu Cheng, Chenyang Si +2
Multi-modal Large Language Models (MLLMs) struggle with long videos due to the need for excessive visual tokens. These tokens exceed massively the context length of MLLMs, resultin…