1 paper · 1 filter
Ryuhei Miyazato, Shunsuke Kitada, Kei Harada
Vision-Language Models (VLMs) excel at multimodal tasks, but they remain vulnerable to hallucinations that are factually incorrect or ungrounded in the input image. Recent work sug…