12 citations · 20 across the 2 of their papers we have counts for
2 papers
cs.CV2021★ 8 cited
Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network
Yehao Li, Yingwei Pan, Ting Yao +2
Despite having impressive vision-language (VL) pretraining with BERT-based encoder for VL understanding, the pretraining of a universal encoder-decoder for both VL understanding an…
cs.CV2019★ 12 cited
Temporal Deformable Convolutional Encoder-Decoder Networks for Video Captioning
Jingwen Chen, Yingwei Pan, Yehao Li +3
It is well believed that video captioning is a fundamental but challenging task in both computer vision and artificial intelligence fields. The prevalent approach is to map an inpu…