26 citations · 61 across the 4 of their papers we have counts for
7 papers
Revealing the Dark Secrets of Masked Image Modeling
Zhenda Xie, Zigang Geng, Jingcheng Hu +3
Masked image modeling (MIM) as pre-training is shown to be effective for numerous vision downstream tasks, but how and where MIM works remain unclear. In this paper, we compare MIM…
iCAR: Bridging Image Classification and Image-text Alignment for Visual Recognition
Yixuan Wei, Yue Cao, Zheng Zhang +4
Image classification, which classifies images by pre-defined categories, has been the dominant approach to visual representation learning over the last decade. Visual learning thro…
Self-Supervised Learning with Swin Transformers
Zhenda Xie, Yutong Lin, Zhuliang Yao +4
We are witnessing a modeling shift from CNN to Transformers in computer vision. In this work, we present a self-supervised learning approach called MoBY, with Vision Transformers a…
Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning
Zhenda Xie, Yutong Lin, Zheng Zhang +3
Contrastive learning methods for unsupervised visual representation learning have reached remarkable levels of transfer performance. We argue that the power of contrastive learning…
Parametric Instance Classification for Unsupervised Visual Feature Learning
Yue Cao, Zhenda Xie, Bin Liu +3
This paper presents parametric instance classification (PIC) for unsupervised visual feature learning. Unlike the state-of-the-art approaches which do instance discrimination in a…
Spatially Adaptive Inference with Stochastic Feature Sampling and Interpolation
Zhenda Xie, Zheng Zhang, Xizhou Zhu +2
In the feature maps of CNNs, there commonly exists considerable spatial redundancy that leads to much repetitive processing. Towards reducing this superfluous computation, we propo…