5 papers · 1 filter
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Chaohong Guo, Xun Mo, Yongwei Nie +3
Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Vision-Language Models (LVLMs) employ Reinforc…
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
Chaohong Guo, Yihan He, Yongwei Nie +3
Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dyn…
Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning
Yihong Huang, Fei Ma, Yihua Shao +4
Vision token pruning has proven to be an effective acceleration technique for the efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent per…
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
Xuyi Yang, Wenhao Zhang, Hongbo Jin +5
Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video…
MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations
Liang Xu, Shaoyang Hua, Zili Lin +6
In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-relat…