1 paper
Yuhan Liu, Lianhui Qin, Shengjie Wang
Large Vision-Language Models (VLMs) have achieved remarkable progress in multimodal understanding, yet they struggle when reasoning over information-intensive images that densely i…