1 paper · 1 filter
Hoang-Bao Le, Aiden Durrant, Thai Son Mai +3
Vision-language models (VLMs) achieve strong cross-modal alignment but remain brittle to negation, often relying on shallow word associations rather than compositional reasoning. F…