1 paper · 1 filter
Biao Tang, Xu Chen, Shuxiang Gou +4
Long-video understanding is constrained by the limited visual input capacity of video multimodal large language models (Video-MLLMs). Existing methods mainly optimize which content…