1 paper · 1 filter
Tao Wu, Li Yang, Gen Zhan +6
Enhancing the temporal understanding of Multimodal Large Language Models (MLLMs) is essential for advancing long-form video analysis, enabling tasks such as temporal localization,…