1 paper
Henghao Zhao, Ge-Peng Ji, Rui Yan +2
The core challenge in video understanding lies in perceiving dynamic content changes over time. However, multimodal large language models struggle with temporal-sensitive video tas…