1 paper
Armaan Sandhu, Abhilasha Senapati, Hima Kammachi
Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whether an interpretability diagnosis of this failure c…