1 paper · 1 filter
Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz +2
Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions…