1 paper · 1 filter
Jinlong Li, Jiaming Ding, Dingfu Lu +8
Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains…