148 citations · 240 across the 27 of their papers we have counts for
Showing cs.MMShow all
2 papers · 1 filter
cs.MM2023★ 1 cited
Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?
Fei Wang, Liang Ding, Jun Rao +3
The multimedia community has shown a significant interest in perceiving and representing the physical world with multimodal pretrained neural network models, and among them, the vi…
cs.MM2022
Dynamic Contrastive Distillation for Image-Text Retrieval
Jun Rao, Liang Ding, Shuhan Qi +4
Although the vision-and-language pretraining (VLP) equipped cross-modal image-text retrieval (ITR) has achieved remarkable progress in the past two years, it suffers from a major d…