7 citations · 11 across the 3 of their papers we have counts for
1 paper · 1 filter
Jun Rao, Liang Ding, Shuhan Qi +4
Although the vision-and-language pretraining (VLP) equipped cross-modal image-text retrieval (ITR) has achieved remarkable progress in the past two years, it suffers from a major d…