1 paper
Jouwon Song, Woohyeong Kim, Kyeongbo Kong
Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that introduces severe latency bo…