1 paper
Alessandro Favero, Luca Zancato, Matthew Trager +5
Generative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phe…