1 paper · 1 filter
Xu Li, Yi Zheng, Yuxuan Liang +5
Large Vision-Language Models (LVLMs) rely on dense visual tokens to capture fine-grained visual information, but processing all these tokens incurs substantial computational and me…