1 paper · 1 filter
Jiyang Tang, Hengyi Li, Yifan Du +1
Although video multimodal large language models (video MLLMs) have achieved substantial progress in video captioning tasks, it remains challenging to adjust the focal emphasis of v…