1 paper
Yutao Jiang, Qiong Wu, Wenhao Lin +2
Recent Multimodal Large Language Models(MLLMs) often use a large number of visual tokens to compensate their visual shortcoming, leading to excessive computation and obvious visual…