11 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 11 cited
VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Xingjian He, Sihan Chen, Fan Ma +7
Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…
cs.CV2023★ 2 cited
CMAE-V: Contrastive Masked Autoencoders for Video Action Recognition
Cheng-Ze Lu, Xiaojie Jin, Zhicheng Huang +3
Contrastive Masked Autoencoder (CMAE), as a new self-supervised framework, has shown its potential of learning expressive feature representations in visual image recognition. This…