2 papers
cs.CV2026
Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding
Baiyang Song, Yuli Lin, Qiong Wu +5
Currently, streaming video understanding is still a daunting task for existing \emph{multimodal large language models} (MLLMs). Its difficulties not only lie in handling the ever-i…
cs.CV2026
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
Tao Chen, Kun Zhang, Qiong Wu +5
Long video understanding is a key challenge that plagues the advancement of \emph{Multimodal Large language Models} (MLLMs). In this paper, we study this problem from the perspecti…