2 citations · 2 across the 11 of their papers we have counts for
1 paper · 1 filter
Yifeng Yao, Yike Yun, Jing Wang +6
Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in image understanding, but long-video are constrained by context windows and computational cost…