#video understanding
7 papers · 1 filter
VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
Haiyue Zhang, Yi Bin, Xun Jiang +5
VisualRouter is a training-free, plug‑and‑play framework that classifies queries as global or local and applies tailored visual sampling strategies to select informative frames, im…
Knowledge-guided Disentanglement with Atomic Actions for Action Recognition
Tianci Wu, Siqi Cao, Guangming Zhu +6
The paper introduces a framework that uses large language models to break down action labels into atomic actions and injects this semantic knowledge into video features to improve…
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…
CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams
Xu Liu, Guikun Chen, Zihao Yan +2
The paper introduces CoVStream, an edge‑cloud system that compresses raw video into compact visual features and captions on the device, sends them to the cloud for graph‑based reas…
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
Dayong Liu, Chao Xu, Weihong Chen +5
The paper introduces CFG-Bench, a benchmark of videos and QA pairs to evaluate how well multimodal language models can generate fine-grained action instructions and higher-order re…
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
Sihan Chen, Jiale Li, Jianghang Lin +1
The paper introduces MoHallBench, a large benchmark designed to evaluate and diagnose motion hallucination—incorrectly inferred human motions—in video large language models, coveri…