1 paper · 1 filter
Zhenkai Wu, Xiaowen Ma, Zhenliang Ni +4
Vision-language models (VLMs) excel at image understanding tasks, but the large number of visual tokens imposes significant computational costs, hindering deployment on mobile devi…