1 paper
Soumyadeep Jana, Pulkit Mittal, Sanasam Ranbir Singh
Large vision-language models (LVLMs) often hallucinate objects that are not present in the input image, largely because visual grounding weakens as decoding progresses. Existing in…