1 paper
Zixiu Ding, Zilin Zhao, Yingjie He +3
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes…