11 citations · 21 across the 2 of their papers we have counts for
1 paper · 1 filter
Fan Ma, Xiaojie Jin, Heng Wang +4
Video-Language Pre-training models have recently significantly improved various multi-modal downstream tasks. Previous dominant works mainly adopt contrastive learning to achieve g…