165 citations
1 paper · 1 filter
Md Jahidul Islam
The adaptation of large-scale Vision-Language Models (VLMs) like CLIP to downstream tasks with extremely limited data -- specifically in the one-shot regime -- is often hindered by…