1 paper
Yichen Guo, Tinghao Wang, Qizhe Zhang +15
Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead,…