VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
arXiv:2105.04906
Abstract
Recent self-supervised methods for image representation learning are based on maximizing the agreement between embedding vectors from different views of the same image. A trivial solution is obtained when the encoder outputs constant vectors. This collapse problem is often avoided through implicit biases in the learning architecture, that often lack a clear justification or interpretation. In this paper, we introduce VICReg (Variance-Invariance-Covariance Regularization), a method that explicitly avoids the collapse problem with a simple regularization term on the variance of the embeddings along each dimension individually. VICReg combines the variance term with a decorrelation mechanism based on redundancy reduction and covariance regularization, and achieves results on par with the state of the art on several downstream tasks. In addition, we show that incorporating our new variance term into other methods helps stabilize the training and leads to performance improvements.
Accepted at ICLR 2022
Cited by in corpus (21)
- Contrastive Masked Autoencoders are Stronger Vision Learners
- Video Transformers: A Survey
- Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
- SoftHebb: Bayesian Inference in Unsupervised Hebbian Soft Winner-Take-All Networks
- Self-Supervised and Invariant Representations for Wireless Localization
- Data Compression and Inference in Cosmology with Self-Supervised Machine Learning
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- DimCL: Dimensional Contrastive Learning For Improving Self-Supervised Learning
- LibAUC: A Deep Learning Library for X-Risk Optimization
- Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
- A Softmax-free Loss Function Based on Predefined Optimal-distribution of Latent Features for Deep Learning Classifier
- MUSE: Music Recommender System with Shuffle Play Recommendation Enhancement
- Hierarchical end-to-end autonomous navigation through few-shot waypoint detection
- KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
- Self-Supervised Representation Learning for Nerve Fiber Distribution Patterns in 3D-PLI
- MACK: Mismodeling Addressed with Contrastive Knowledge
- Unsupervised End-to-End Training with a Self-Defined Target
- Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets?
- Kilonova Light Curve Parameter Estimation Using Likelihood-Free Inference
- SERE: Exploring Feature Self-relation for Self-supervised Transformer
- Data-Driven Self-Supervised Graph Representation Learning