1 paper
Jun Ling, Tao Huang, Junzhuo Liu +2
Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens that substantially increase downstream la…