collaborators

6 papers

cs.CV2026

QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

Jun Peng, Baiyang Song, Jie Li +4

Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This chal…

cs.CV2026

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding

Baiyang Song, Yuli Lin, Qiong Wu +5

Currently, streaming video understanding is still a daunting task for existing \emph{multimodal large language models} (MLLMs). Its difficulties not only lie in handling the ever-i…

cs.CV2026

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling

Kun Zhang, Chenxin Fang, Tao Chen +4

Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint. Thus, keyframe selection i…

cs.CV2026

KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs

Baiyang Song, Jun Peng, Yuxin Zhang +3

Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static fra…

cs.CV2025

Omni-Referring Image Segmentation

Qiancheng Zheng, Yunhang Shen, Gen Luo +5

In this paper, we propose a novel task termed Omni-Referring Image Segmentation (OmniRIS) towards highly generalized image segmentation. Compared with existing unimodally condition…

cs.CV2025

Grounded Chain-of-Thought for Multimodal Large Language Models

Qiong Wu, Xiangcong Yang, Yiyi Zhou +4

Despite great progress, existing multimodal large language models (MLLMs) are prone to visual hallucination, greatly impeding their trustworthy applications. In this paper, we stud…