1 paper
Biao Tang, Xu Chen, Shuxiang Gou +3
Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames and are therefore expensive…