1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Suyang Xi, Chenxi Yang, Hong Ding +4
Multimodal large language models (MLLMs) often fail in fine-grained visual question answering, producing hallucinations about object identities, positions, and relations because te…