35 citations · 35 across the 1 of their papers we have counts for
1 paper
Peng Jin, Jinfa Huang, Fenglin Liu +5
Most video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to t…