70 citations · 241 across the 13 of their papers we have counts for
19 papers
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
Yuxin Fang, Wen Wang, Binhui Xie +6
We launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data. EVA is a vanilla ViT pre-trained to re…
Could Giant Pretrained Image Models Extract Universal Representations?
Yutong Lin, Ze Liu, Zheng Zhang +4
Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few pa…
Revealing the Dark Secrets of Masked Image Modeling
Zhenda Xie, Zigang Geng, Jingcheng Hu +3
Masked image modeling (MIM) as pre-training is shown to be effective for numerous vision downstream tasks, but how and where MIM works remain unclear. In this paper, we compare MIM…
iCAR: Bridging Image Classification and Image-text Alignment for Visual Recognition
Yixuan Wei, Yue Cao, Zheng Zhang +4
Image classification, which classifies images by pre-defined categories, has been the dominant approach to visual representation learning over the last decade. Visual learning thro…
Bootstrap Your Object Detector via Mixed Training
Mengde Xu, Zheng Zhang, Fangyun Wei +5
We introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by ut…
Video Swin Transformer
Ze Liu, Jia Ning, Yue Cao +4
The vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchm…