1 paper · 1 filter
Chengyan Wang, Hanliang Xie, Yueyi Yang +1
While Multimodal Large Language Models (MLLMs) excel in general video understanding, their capability in fine-grained and motion-centric tasks remains limited. This limitation is p…