1 paper
Tianyu Fu, Tengxuan Liu, Qinghao Han +5
The increasing demand to process long and high-resolution videos significantly burdens Large Vision-Language Models (LVLMs) due to the enormous number of visual tokens. Existing to…