1 paper
Deniz Kizaroğlu, Ülku Tuncer Küçüktas, Emre Çakmakyurdu +1
Few-shot adaptation of vision-language models (VLMs) like CLIP typically relies on learning textual prompts matched to global image embeddings. Recent works extend this paradigm by…