1 paper
Jungbeom Lee, Sanghyuk Chun, Sangdoo Yun
Recent Vision-Language Pre-training (VLP) models have demonstrated significant advancements. Nevertheless, these models heavily rely on image-text pairs that capture only coarse an…