61 citations · 156 across the 14 of their papers we have counts for
1 paper · 1 filter
Shihong Liu, Zhiqiu Lin, Samuel Yu +4
Vision-language models (VLMs) pre-trained on web-scale datasets have demonstrated remarkable capabilities on downstream tasks when fine-tuned with minimal data. However, many VLMs…