1 paper · 1 filter
Dylan Sam, Devin Willmott, Joao D. Semedo +1
Vision-language models (VLMs) such as CLIP are trained via contrastive learning between text and image pairs, resulting in aligned image and text embeddings that are useful for man…