1 paper
Jinchang Zhu, Rong Fu, Yi Ding +3
Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image is compressed into a low-resolu…