1 paper
Walid Bousselham, Sofian Chaybouti, Christian Rupprecht +2
Vision-language foundation models such as CLIP have achieved tremendous results in global vision-language alignment, but still show some limitations in creating representations for…