2 papers
cs.LG2025
Finetuning CLIP to Reason about Pairwise Differences
Dylan Sam, Devin Willmott, Joao D. Semedo +1
Vision-language models (VLMs) such as CLIP are trained via contrastive learning between text and image pairs, resulting in aligned image and text embeddings that are useful for man…
cs.CV2024
HyperCLIP: Adapting Vision-Language models with Hypernetworks
Victor Akinwande, Mohammad Sadegh Norouzzadeh, Devin Willmott +3
Self-supervised vision-language models trained with contrastive objectives form the basis of current state-of-the-art methods in AI vision tasks. The success of these models is a d…