35 citations · 113 across the 10 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022★ 35 cited
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations
Peng Jin, Jinfa Huang, Fenglin Liu +5
Most video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to t…
cs.CV2022★ 23 cited
How to Understand Masked Autoencoders
Shuhao Cao, Peng Xu, David A. Clifton
"Masked Autoencoders (MAE) Are Scalable Vision Learners" revolutionizes the self-supervised learning method in that it not only achieves the state-of-the-art for image pre-training…