1 paper
Haiji Liang, Pengfei Zhou, Zhenglin Wan +3
Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existing vision token pruning method…