1 paper
Aditi Gupta, Yossi Gandelsman
We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and thr…