1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Avik Pal, Max van Spengler, Guido Maria D'Amely di Melendugno +3
Image-text representation learning forms a cornerstone in vision-language models, where pairs of images and textual descriptions are contrastively aligned in a shared embedding spa…