collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Visual Token Coding for Video Multimodal Large Language Models

Chenxin Fang, Tao Chen, JunChao You +3

In this paper, we propose a new token compression paradigm for video Multimodal Large Language Models (MLLMs), termed Visual Token Coding (VTC). Inspired by classical video coding…

cs.CV2026

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

Jun Peng, Baiyang Song, Jie Li +4

Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This chal…

cs.CV2026

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding

Baiyang Song, Yuli Lin, Qiong Wu +5

Currently, streaming video understanding is still a daunting task for existing \emph{multimodal large language models} (MLLMs). Its difficulties not only lie in handling the ever-i…

cs.CV2025

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

Tao Chen, Shaobo Ju, Qiong Wu +6

Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose an effective and efficient para…

cs.CV2025

AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection

Shuheng Zhang, Yuqi Liu, Hongbo Zhou +4

Despite great progress, text-driven long video editing is still notoriously challenging mainly due to excessive memory overhead. Although recent efforts have simplified this task i…