1 paper · 1 filter
Jiaying Zhu, Yurui Zhu, Xin Lu +5
Multimodal Large Language Models (MLLMs) encounter significant computational and memory bottlenecks from the massive number of visual tokens generated by high-resolution images or…