1 paper
Guangyu Sun, Archit Singhal, Burak Uzkent +3
Video Large Language Models (VLMs) have achieved strong performance on various vision-language tasks, yet their practical use is limited by the massive number of visual tokens prod…