1 paper
Qiankun Ma, Ziyao Zhang, Haofei Wang +3
Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead an…