50 citations · 199 across the 15 of their papers we have counts for
17 papers
X-modaler: A Versatile and High-performance Codebase for Cross-modal Analytics
Yehao Li, Yingwei Pan, Jingwen Chen +2
With the rise and development of deep learning over the past decade, there has been a steady momentum of innovation and breakthroughs that convincingly push the state-of-the-art of…
Contextual Transformer Networks for Visual Recognition
Yehao Li, Ting Yao, Yingwei Pan +1
Transformer with self-attention has led to the revolutionizing of natural language processing field, and recently inspires the emergence of Transformer-style architecture design wi…
Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder Network
Yehao Li, Yingwei Pan, Ting Yao +2
Despite having impressive vision-language (VL) pretraining with BERT-based encoder for VL understanding, the pretraining of a universal encoder-decoder for both VL understanding an…
Pre-training for Video Captioning Challenge 2020 Summary
Yingwei Pan, Jun Xu, Yehao Li +2
The Pre-training for Video Captioning Challenge 2020 Summary: results and challenge participants' technical reports.
Auto-captions on GIF: A Large-scale Video-sentence Dataset for Vision-language Pre-training
Yingwei Pan, Yehao Li, Jianjie Luo +3
In this work, we present Auto-captions on GIF, which is a new large-scale pre-training dataset for generic video understanding. All video-sentence pairs are created by automaticall…
Exploring Category-Agnostic Clusters for Open-Set Domain Adaptation
Yingwei Pan, Ting Yao, Yehao Li +2
Unsupervised domain adaptation has received significant attention in recent years. Most of existing works tackle the closed-set scenario, assuming that the source and target domain…