32 citations · 35 across the 6 of their papers we have counts for
1 paper · 1 filter
Simon Schrodi, David T. Hoffmann, Max Argus +2
Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, l…