1 paper
Fengze Yang, Bo Yu, Xuewen Luo +2
Vision-Language Models (VLMs) face severe memory and latency bottlenecks due to high-resolution visual tokens. While current token reduction methods theoretically save FLOPs, post-…