1 paper
Amita Kamath, Cheng-Yu Hsieh, Kai-Wei Chang +1
Several benchmarks have concluded that our best vision-language models (e.g., CLIP) are lacking in compositionality. Given an image, these benchmarks probe a model's ability to ide…