Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
arXiv:2005.10242
Abstract
Contrastive representation learning has been outstandingly successful in practice. In this work, we identify two key properties related to the contrastive loss: (1) alignment (closeness) of features from positive pairs, and (2) uniformity of the induced distribution of the (normalized) features on the hypersphere. We prove that, asymptotically, the contrastive loss optimizes these properties, and analyze their positive effects on downstream tasks. Empirically, we introduce an optimizable metric to quantify each property. Extensive experiments on standard vision and language datasets confirm the strong agreement between both metrics and downstream task performance. Remarkably, directly optimizing for these two metrics leads to representations with comparable or better performance at downstream tasks than contrastive learning. Project Page: https://tongzhouwang.info/hypersphere Code: https://github.com/SsnL/align_uniform , https://github.com/SsnL/moco_align_uniform
International Conference on Machine Learning (ICML), 2020
References in corpus (7)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- A Simple Framework for Contrastive Learning of Visual Representations
- Improved Baselines with Momentum Contrastive Learning
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- A Theoretical Analysis of Contrastive Unsupervised Representation Learning
- von Mises-Fisher Mixture Model-based Deep learning: Application to Face Verification
Cited by in corpus (28)
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Training Vision Transformers for Image Retrieval
- Supervised Contrastive Learning
- Learning and Evaluating Representations for Deep One-class Classification
- SEED: Self-supervised Distillation For Visual Representation
- Intriguing Properties of Contrastive Losses
- Predicting What You Already Know Helps: Provable Self-Supervised Learning
- ReSSL: Relational Self-Supervised Learning with Weak Augmentation
- Multi-Label Contrastive Learning for Abstract Visual Reasoning
- Understanding Self-supervised Learning with Dual Deep Networks
- Extracting Sentence Embeddings from Pretrained Transformer Models
- Understanding the Behaviour of Contrastive Loss
- Limits on Inferring T-cell Specificity from Partial Information
- Debiased Contrastive Learning
- Toward a Better Understanding of Loss Functions for Collaborative Filtering
- Run Away From your Teacher: Understanding BYOL by a Novel Self-Supervised Approach
- Improving Transformation Invariance in Contrastive Representation Learning
- Evolution Is All You Need: Phylogenetic Augmentation for Contrastive Learning
- Center-wise Local Image Mixture For Contrastive Representation Learning
- EqCo: Equivalent Rules for Self-supervised Contrastive Learning
- Function Contrastive Learning of Transferable Meta-Representations
- S2SD: Simultaneous Similarity-based Self-Distillation for Deep Metric Learning
- Disentangled Contrastive Learning for Learning Robust Textual Representations
- Uniform Priors for Data-Efficient Transfer
- Complementary Relation Contrastive Distillation
- Momentum Contrastive Autoencoder: Using Contrastive Learning for Latent Space Distribution Matching in WAE
- Understanding the Role of Self-Supervised Learning in Out-of-Distribution Detection Task
- Contrastive learning of strong-mixing continuous-time stochastic processes