6 papers
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
Jun Peng, Baiyang Song, Jie Li +4
Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This chal…
Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding
Baiyang Song, Yuli Lin, Qiong Wu +5
Currently, streaming video understanding is still a daunting task for existing \emph{multimodal large language models} (MLLMs). Its difficulties not only lie in handling the ever-i…
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling
Kun Zhang, Chenxin Fang, Tao Chen +4
Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint. Thus, keyframe selection i…
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
Baiyang Song, Jun Peng, Yuxin Zhang +3
Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static fra…
Omni-Referring Image Segmentation
Qiancheng Zheng, Yunhang Shen, Gen Luo +5
In this paper, we propose a novel task termed Omni-Referring Image Segmentation (OmniRIS) towards highly generalized image segmentation. Compared with existing unimodally condition…
Grounded Chain-of-Thought for Multimodal Large Language Models
Qiong Wu, Xiangcong Yang, Yiyi Zhou +4
Despite great progress, existing multimodal large language models (MLLMs) are prone to visual hallucination, greatly impeding their trustworthy applications. In this paper, we stud…