1 paper
Ameen Ali, Tamim Zoabi, Lidor Brami +1
Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a phenomenon known as object…