342 citations · 689 across the 28 of their papers we have counts for
7 papers · 2 filters
BEVT: BERT Pretraining of Video Transformers
Rui Wang, Dongdong Chen, Zuxuan Wu +6
This paper studies the BERT pretraining of video transformers. It is a straightforward but worth-studying extension given the recent success from BERT pretraining of image transfor…
Florence: A New Foundation Model for Computer Vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen +20
Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human visio…
MicroNet: Improving Image Recognition with Extremely Low FLOPs
Yunsheng Li, Yinpeng Chen, Xiyang Dai +6
This paper aims at addressing the problem of substantial performance degradation at extremely low computational cost (e.g. 5M FLOPs on ImageNet classification). We found that two f…
Mobile-Former: Bridging MobileNet and Transformer
Yinpeng Chen, Xiyang Dai, Dongdong Chen +4
We present Mobile-Former, a parallel design of MobileNet and transformer with a two-way bridge in between. This structure leverages the advantages of MobileNet at local processing…
Dynamic Head: Unifying Object Detection Heads with Attentions
Xiyang Dai, Yinpeng Chen, Bin Xiao +4
The complex nature of combining localization and classification in object detection has resulted in the flourished development of methods. Previous works tried to improve the perfo…
CvT: Introducing Convolutions to Vision Transformers
Haiping Wu, Bin Xiao, Noel Codella +4
We present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convo…