32 citations · 39 across the 9 of their papers we have counts for
1 paper · 2 filters
Simon Schrodi, David T. Hoffmann, Max Argus +2
Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, l…