1 paper · 1 filter
Yi Wang, Yinan He, Yizhuo Li +13
This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understand…