12 citations · 16 across the 3 of their papers we have counts for
1 paper · 1 filter
Alex Jinpeng Wang, Yixiao Ge, Rui Yan +7
Mainstream Video-Language Pre-training models \cite{actbert,clipbert,violet} consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They purs…