7 citations · 7 across the 1 of their papers we have counts for
1 paper · 1 filter
Xilun Chen, Lili Yu, Wenhan Xiong +3
We propose a new two-stage pre-training framework for video-to-text generation tasks such as video captioning and video question answering: A generative encoder-decoder model is fi…