1 paper
Junshan Hu, Jialiang Mao, Zhikang Liu +3
Conventional Vision-Language Models(VLMs) typically utilize a fixed number of vision tokens, regardless of task complexity. This one-size-fits-all strategy introduces notable ineff…