1 paper
Qizhe Zhang, Aosong Cheng, Ming Lu +6
Large vision-language models (LVLMs) generally contain significantly more visual tokens than their textual counterparts, resulting in a considerable computational burden. Recent ef…