4 papers
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling
Kun Zhang, Chenxin Fang, Tao Chen +4
Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint. Thus, keyframe selection i…
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
Tao Chen, Shaobo Ju, Qiong Wu +6
Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose an effective and efficient para…
CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation
Yuhang Ming, Chenxin Fang, Xingyuan Yu +4
Recent advances in Gaussian Splatting based 3D scene representation have shown two major trends: semantics-oriented approaches that focus on high-level understanding but lack expli…
Grounded Chain-of-Thought for Multimodal Large Language Models
Qiong Wu, Xiangcong Yang, Yiyi Zhou +4
Despite great progress, existing multimodal large language models (MLLMs) are prone to visual hallucination, greatly impeding their trustworthy applications. In this paper, we stud…