1 paper
Yuke Zhu, Chi Xie, Shuang Liang +2
Recent advances on Multi-modal Large Language Models have demonstrated that high-resolution image input is crucial for model capabilities, especially for fine-grained tasks. Howeve…