1 paper
Fanrui Zhang, Jiawei Liu, Jiaying Zhu +4
Multimodal Large Language Models (MLLMs), such as GPT4o, have shown strong capabilities in visual reasoning and explanation generation. However, despite these strengths, they face…