10 citations · 10 across the 1 of their papers we have counts for
1 paper
Yue Ma, Tianyu Yang, Yin Shan +1
This paper presents SimVTP: a Simple Video-Text Pretraining framework via masked autoencoders. We randomly mask out the spatial-temporal tubes of input video and the word tokens of…