1 paper
Biao Tang, Xu Chen, Shuxiang Gou +4
Long-video understanding is constrained by the limited visual input capacity of video multimodal large language models (Video-MLLMs). Existing methods mainly optimize which content…