26 citations · 77 across the 5 of their papers we have counts for
1 paper · 1 filter
Yichen Guo, Tinghao Wang, Qizhe Zhang +15
Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead,…