From the 1 of 41 linked papers with an AI index.
1 paper · 1 filter
Huanxuan Liao, Zhongtao Jiang, Yupu Hao +6
Multimodal Large Language Models (MLLMs) achieve stronger visual understanding by scaling input fidelity, yet the resulting visual token growth makes jointly sustaining high spatia…