9 citations · 10 across the 4 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022★ 9 cited
VLMAE: Vision-Language Masked Autoencoder
Sunan He, Taian Guo, Tao Dai +4
Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data…
cs.CV2022
Exploiting Feature Diversity for Make-up Temporal Video Grounding
Xiujun Shu, Wei Wen, Taian Guo +3
This technical report presents the 3rd winning solution for MTVG, a new task introduced in the 4-th Person in Context (PIC) Challenge at ACM MM 2022. MTVG aims at localizing the te…