15 citations · 15 across the 5 of their papers we have counts for
1 paper · 2 filters
Jonghyun Song, Youngjune Lee, Gyu-Hwung Cho +3
Vision-Language Pretrained (VLP) models have achieved impressive performance on multimodal tasks, including text-image retrieval, based on dense representations. Meanwhile, Learned…