1 paper
Chenglin Li, Qianglong Chen, fengtao +1
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding tasks. However, they continue to struggle with long-form videos because of an ineffici…