1 paper · 1 filter
Mengjie Zhang, Qihui Zhu, Tao Zhang +10
Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal v…