1 paper
Jun Rao, Liang Ding, Shuhan Qi +4
Although the vision-and-language pretraining (VLP) equipped cross-modal image-text retrieval (ITR) has achieved remarkable progress in the past two years, it suffers from a major d…