26 citations · 26 across the 1 of their papers we have counts for
1 paper
Jiarui Yu, Haoran Li, Yanbin Hao +3
The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated…