1 paper
Junyang Zhang, Puhan Luo, Chen Tang +2
Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. However, processing the large number of frames in a video incurs substan…