1 paper
Dong-Jae Lee, Sunghyun Baek, Junmo Kim
Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens. Ex…