219 citations · 553 across the 7 of their papers we have counts for
12 papers · 1 filter
Exploring Long-Sequence Masked Autoencoders
Ronghang Hu, Shoubhik Debnath, Saining Xie +1
Masked Autoencoding (MAE) has emerged as an effective approach for pre-training representations across multiple domains. In contrast to discrete tokens in natural languages, the in…
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu +3
The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification mo…
A Fistful of Words: Learning Transferable Visual Models from Bag-of-Words Supervision
Ajinkya Tejankar, Maziar Sanjabi, Bichen Wu +4
Using natural language as a supervision for training visual recognition models holds great promise. Recent works have shown that if such supervision is used in the form of alignmen…
An Empirical Study of Training Self-Supervised Vision Transformers
Xinlei Chen, Saining Xie, Kaiming He
This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know baseline given the recent progress in computer vision: self-supervise…
Exploring Data-Efficient 3D Scene Understanding with Contrastive Scene Contexts
Ji Hou, Benjamin Graham, Matthias Nießner +1
The rapid progress in 3D scene understanding has come with growing demand for data; however, collecting and annotating 3D scenes (e.g. point clouds) are notoriously hard. For examp…
PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding
Saining Xie, Jiatao Gu, Demi Guo +3
Arguably one of the top success stories of deep learning is transfer learning. The finding that pre-training a network on a rich source set (eg., ImageNet) can help boost performan…