20 citations · 20 across the 1 of their papers we have counts for
1 paper
Shen Yan, Tao Zhu, Zirui Wang +5
We explore an efficient approach to establish a foundational video-text model. We present VideoCoCa that maximally reuses a pretrained image-text contrastive captioner (CoCa) model…