7 papers · 1 filter
Conditional Multi-Event Temporal Grounding in Long-Form Video
Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez +12
Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositi…
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
Yiming Liang, Yixiao Chen, Yiyang Zhou +8
Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-language models (LVLMs) often a…
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
Andong Deng, Dawei Du, Zhenfang Chen +7
Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. Whi…
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
Andong Deng, Taojiannan Yang, Shoubin Yu +5
Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and ch…
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
Yuanhao Zou, Shengji Jin, Andong Deng +3
Effectively applying Vision-Language Models (VLMs) to Video Question Answering (VideoQA) hinges on selecting a concise yet comprehensive set of frames, as processing entire videos…
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
Andong Deng, Tongjia Chen, Shoubin Yu +6
In this paper, we introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the inp…