1 paper
Yixin Wan, Tianle Zheng, Kai-Wei Chang
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative…