1 paper
Guo Chen, Yicheng Liu, Yifei Huang +6
Most existing video understanding benchmarks for multimodal large language models (MLLMs) focus only on short videos. The limited number of benchmarks for long video understanding…