1 paper · 1 filter
Zhihui Guo, Xin Man, Hui Xu +4
Multimodal Large Language Models (MLLMs) excel in vision-language tasks such as image captioning but remain prone to object hallucinations, where they describe objects that do not…