1 paper · 1 filter
Amartya Bhattacharya
Vision-language models (VLMs) excel at image-text retrieval yet persistently fail at compositional reasoning, distinguishing captions that share the same words but differ in relati…