1 paper
Zheyu Zhang, Ziqi Pang, Shixing Chen +3
Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens…