1 paper
Yifeng Yao, Yike Yun, Jing Wang +6
Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in image understanding, but long-video are constrained by context windows and computational cost…