1 paper
Chengyan Wang, Hanliang Xie, Yueyi Yang +1
While Multimodal Large Language Models (MLLMs) excel in general video understanding, their capability in fine-grained and motion-centric tasks remains limited. This limitation is p…