1 paper · 1 filter
Hongchen Wei, Zhihong Tan, Yaosi Hu +2
Large Multimodal Models (LMMs) have demonstrated exceptional performance in video captioning tasks, particularly for short videos. However, as the length of the video increases, ge…