1 paper · 1 filter
Xuan Yu, Dayan Guan, Yanfeng Gu
Multimodal Large Language Models (MLLM) often struggle to interpret high-resolution images accurately, where fine-grained details are crucial for complex visual understanding. We i…