1 paper
Peizheng Yan, Yu Zhao, Liang Xie +3
Large vision-language models perform well on short- and medium-length video understanding but still struggle to maintain coherent event memory and recover long-range relationships…