9 citations · 18 across the 3 of their papers we have counts for
3 papers
cs.CV2022★ 9 cited
See Finer, See More: Implicit Modality Alignment for Text-based Person Retrieval
Xiujun Shu, Wei Wen, Haoqian Wu +5
Text-based person retrieval aims to find the query person based on a textual description. The key is to learn a common latent space mapping between visual-textual modalities. To ac…
cs.CV2022★ 9 cited
VLMAE: Vision-Language Masked Autoencoder
Sunan He, Taian Guo, Tao Dai +4
Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data…
cs.CV2022
Exploiting Feature Diversity for Make-up Temporal Video Grounding
Xiujun Shu, Wei Wen, Taian Guo +3
This technical report presents the 3rd winning solution for MTVG, a new task introduced in the 4-th Person in Context (PIC) Challenge at ACM MM 2022. MTVG aims at localizing the te…