Bootstrap your own latent: A new approach to self-supervised Learning
arXiv:2006.07733
Abstract
We introduce Bootstrap Your Own Latent (BYOL), a new approach to self-supervised image representation learning. BYOL relies on two neural networks, referred to as online and target networks, that interact and learn from each other. From an augmented view of an image, we train the online network to predict the target network representation of the same image under a different augmented view. At the same time, we update the target network with a slow-moving average of the online network. While state-of-the art methods rely on negative pairs, BYOL achieves a new state of the art without them. BYOL reaches top-1 classification accuracy on ImageNet using a linear evaluation with a ResNet-50 architecture and with a larger ResNet. We show that BYOL performs on par or better than the current state of the art on both transfer and semi-supervised benchmarks. Our implementation and pretrained models are given on GitHub.
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- On the difficulty of training Recurrent Neural Networks
- Semi-Supervised Learning with Deep Generative Models
- Temporal Ensembling for Semi-Supervised Learning
- Large Batch Training of Convolutional Networks
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Learning with Pseudo-Ensembles
- A Theoretical Analysis of Contrastive Unsupervised Representation Learning
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- MaxUp: A Simple Way to Improve Generalization of Neural Network Training
Cited by in corpus (7)
- MoPro: Webly Supervised Learning with Momentum Prototypes
- Conditional Negative Sampling for Contrastive Learning of Visual Representations
- Hybrid Discriminative-Generative Training via Contrastive Learning
- False Detection (Positives and Negatives) in Object Detection
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- A Simple Framework for Uncertainty in Contrastive Learning
- Langevin Cooling for Domain Translation