1 paper
Zeshang Li, Shuoyang Zhang
Vision-Language Models (VLMs) hallucinate objects that are not present, and a growing line of work tries to curb this by feeding the model its own generated caption as auxiliary ev…