1 paper · 1 filter
Dhananjay Ashok, Ashutosh Chaubey, Hirona J. Arai +2
Through a controlled study, we identify a systematic deficiency in the multimodal grounding of Vision Language Models (VLMs). While VLMs can recall factual associations when provid…