3 papers
cs.CV2026
Visual Token Coding for Video Multimodal Large Language Models
Chenxin Fang, Tao Chen, JunChao You +3
In this paper, we propose a new token compression paradigm for video Multimodal Large Language Models (MLLMs), termed Visual Token Coding (VTC). Inspired by classical video coding…
cs.CV2026
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
Jun Peng, Baiyang Song, Jie Li +4
Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This chal…
cs.CV2025
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
Tao Chen, Shaobo Ju, Qiong Wu +6
Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose an effective and efficient para…