1 paper · 1 filter
Linhao Yu, Xinguang Ji, Yahui Liu +7
Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, existing benchmarks and evaluation protocols suffe…