EqCo: Equivalent Rules for Self-supervised Contrastive Learning
arXiv:2010.01929
Abstract
In this paper, we propose EqCo (Equivalent Rules for Contrastive Learning) to make self-supervised learning irrelevant to the number of negative samples in the contrastive learning framework. Inspired by the InfoMax principle, we point that the margin term in contrastive loss needs to be adaptively scaled according to the number of negative pairs in order to keep steady mutual information bound and gradient magnitude. EqCo bridges the performance gap among a wide range of negative sample sizes, so that for the first time, we can use only a few negative pairs (e.g., 16 per query) to perform self-supervised contrastive training on large-scale vision datasets like ImageNet, while with almost no accuracy drop. This is quite a contrast to the widely used large batch training or memory bank mechanism in current practices. Equipped with EqCo, our simplified MoCo (SiMo) achieves comparable accuracy with MoCov2 on ImageNet (linear evaluation protocol) while only involves 16 negative pairs per query instead of 65536, suggesting that large quantities of negative samples is not a critical factor in contrastive learning frameworks.
References in corpus (19)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- A Simple Framework for Contrastive Learning of Visual Representations
- Bootstrap your own latent: A new approach to self-supervised Learning
- Improved Baselines with Momentum Contrastive Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- One weird trick for parallelizing convolutional neural networks
- Learning Representations by Maximizing Mutual Information Across Views
- What Makes for Good Views for Contrastive Learning?
- Large Batch Training of Convolutional Networks
- Big Self-Supervised Models are Strong Semi-Supervised Learners
- Exploring Simple Siamese Representation Learning
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
- Circle Loss: A Unified Perspective of Pair Similarity Optimization
- Local Aggregation for Unsupervised Learning of Visual Embeddings
- Whitening for Self-Supervised Representation Learning
- Parametric Instance Classification for Unsupervised Visual Feature Learning
- BYOL works even without batch statistics
- Debiased Contrastive Learning
- Delving into Inter-Image Invariance for Unsupervised Visual Representations