2 papers
cs.CV2024
Updating CLIP to Prefer Descriptions Over Captions
Amir Zur, Elisa Kreiss, Karel D'Oosterlinck +2
Although CLIPScore is a powerful generic metric that captures the similarity between a text and an image, it fails to distinguish between a caption that is meant to complement the…
cs.CL2024
CommVQA: Situating Visual Question Answering in Communicative Contexts
Nandita Shankar Naik, Christopher Potts, Elisa Kreiss
Current visual question answering (VQA) models tend to be trained and evaluated on image-question pairs in isolation. However, the questions people ask are dependent on their infor…