1 citations · 1 across the 1 of their papers we have counts for
1 paper
Suyang Xi, Chenxi Yang, Hong Ding +4
Multimodal large language models (MLLMs) often fail in fine-grained visual question answering, producing hallucinations about object identities, positions, and relations because te…