17 papers
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
Jun Peng, Baiyang Song, Jie Li +4
Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This chal…
Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding
Baiyang Song, Yuli Lin, Qiong Wu +5
Currently, streaming video understanding is still a daunting task for existing \emph{multimodal large language models} (MLLMs). Its difficulties not only lie in handling the ever-i…
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling
Kun Zhang, Chenxin Fang, Tao Chen +4
Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint. Thus, keyframe selection i…
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
Shaobo Ju, Baiyang Song, Tao Chen +6
Due to the great saving of computation and memory overhead, token compression has become a research hot-spot for MLLMs and achieved remarkable progress in image-language tasks. How…
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
Tao Chen, Shaobo Ju, Qiong Wu +6
Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose an effective and efficient para…
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
Tao Chen, Kun Zhang, Qiong Wu +5
Long video understanding is a key challenge that plagues the advancement of \emph{Multimodal Large language Models} (MLLMs). In this paper, we study this problem from the perspecti…