1 paper
Hongyu Lu, Feng Zhang, Wenwei Jin +5
Large vision-language models (LVLMs) achieve strong multimodal understanding, but their inference cost grows rapidly with the number of visual tokens, especially for high-resolutio…