1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Jiyang Tang, Hengyi Li, Yifan Du +1
Although video multimodal large language models (video MLLMs) have achieved substantial progress in video captioning tasks, it remains challenging to adjust the focal emphasis of v…