1 paper
Xinyu Huang, Yuhao Dong, Weiwei Tian +3
State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous visual tokens, many of which ar…