1 paper · 1 filter
Jonghyun Song, Youngjune Lee, Gyu-Hwung Cho +3
Vision-Language Pretrained (VLP) models have achieved impressive performance on multimodal tasks, including text-image retrieval, based on dense representations. Meanwhile, Learned…