7 citations · 17 across the 4 of their papers we have counts for
4 papers
UATVR: Uncertainty-Adaptive Text-Video Retrieval
Bo Fang, Wenhao Wu, Chang Liu +6
With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted…
It Takes Two: Masked Appearance-Motion Modeling for Self-supervised Video Transformer Pre-training
Yuxin Song, Min Yang, Wenhao Wu +3
Self-supervised video transformer pre-training has recently benefited from the mask-and-predict pipeline. They have demonstrated outstanding effectiveness on downstream video tasks…
CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
Haoran Wang, Dongliang He, Wenhao Wu +7
Image-Text Retrieval (ITR) is challenging in bridging visual and lingual modalities. Contrastive learning has been adopted by most prior arts. Except for limited amount of negative…
DALG: Deep Attentive Local and Global Modeling for Image Retrieval
Yuxin Song, Ruolin Zhu, Min Yang +1
Deeply learned representations have achieved superior image retrieval performance in a retrieve-then-rerank manner. Recent state-of-the-art single stage model, which heuristically…