39 citations · 51 across the 5 of their papers we have counts for
4 papers · 1 filter
Dual Vision Transformer
Ting Yao, Yehao Li, Yingwei Pan +3
Prior works have proposed several strategies to reduce the computational cost of self-attention mechanism. Many of these works consider decomposing the self-attention procedure int…
Wave-ViT: Unifying Wavelet and Transformers for Visual Representation Learning
Ting Yao, Yingwei Pan, Yehao Li +2
Multi-scale Vision Transformer (ViT) has emerged as a powerful backbone for computer vision tasks, while the self-attention computation in Transformer scales quadratically w.r.t. t…
CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising
Jianjie Luo, Yehao Li, Yingwei Pan +3
BERT-type structure has led to the revolution of vision-language pre-training and the achievement of state-of-the-art results on numerous vision-language downstream tasks. Existing…
Boosting Image Captioning with Attributes
Ting Yao, Yingwei Pan, Yehao Li +2
Automatically describing an image with a natural language has been an emerging challenge in both fields of computer vision and natural language processing. In this paper, we presen…