11 citations · 21 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 11 cited
VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Xingjian He, Sihan Chen, Fan Ma +7
Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…
cs.CV2023★ 10 cited
Temporal Perceiving Video-Language Pre-training
Fan Ma, Xiaojie Jin, Heng Wang +4
Video-Language Pre-training models have recently significantly improved various multi-modal downstream tasks. Previous dominant works mainly adopt contrastive learning to achieve g…