1 paper
Zhaoyang Wei, Zipeng Wang, Yushe Cao +10
Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluatio…