1 paper
Yibo Peng, Peng Xia, Ding Zhong +6
Despite the rapid advancements in Multimodal Large Language Models (MLLMs), a critical question regarding their visual grounding mechanism remains unanswered: do these models genui…