6 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Jiayi Lin, Shaogang Gong
A vision-language foundation model pretrained on very large-scale image-text paired data has the potential to provide generalizable knowledge representation for downstream visual r…