1 paper
Wenhao Xu, Xin Dong, Yue Li +2
Video large language models have demonstrated strong video understanding capabilities but suffer from high inference costs due to the massive number of tokens in long videos. Inspi…