42 citations · 80 across the 3 of their papers we have counts for
1 paper · 1 filter
Hu Xu, Gargi Ghosh, Po-Yao Huang +5
We present VideoCLIP, a contrastive approach to pre-train a unified model for zero-shot video and text understanding, without using any labels on downstream tasks. VideoCLIP trains…