A Unified Mixture-View Framework for Unsupervised Representation Learning
arXiv:2011.13356
Abstract
Recent unsupervised contrastive representation learning follows a Single Instance Multi-view (SIM) paradigm where positive pairs are usually constructed with intra-image data augmentation. In this paper, we propose an effective approach called Beyond Single Instance Multi-view (BSIM). Specifically, we impose more accurate instance discrimination capability by measuring the joint similarity between two randomly sampled instances and their mixture, namely spurious-positive pairs. We believe that learning joint similarity helps to improve the performance when encoded features are distributed more evenly in the latent space. We apply it as an orthogonal improvement for unsupervised contrastive representation learning, including current outstanding methods SimCLR, MoCo, and BYOL. We evaluate our learned representations on many downstream benchmarks like linear classification on ImageNet-1k and PASCAL VOC 2007, object detection on MS COCO 2017 and VOC, etc. We obtain substantial gains with a large margin almost on all these tasks compared with prior arts.
BMVC 2022
References in corpus (12)
- A Simple Framework for Contrastive Learning of Visual Representations
- Bootstrap your own latent: A new approach to self-supervised Learning
- Improved Regularization of Convolutional Neural Networks with Cutout
- Improved Baselines with Momentum Contrastive Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- What Makes for Good Views for Contrastive Learning?
- Large Batch Training of Convolutional Networks
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Local Aggregation for Unsupervised Learning of Visual Embeddings
- Affinity and Diversity: Quantifying Mechanisms of Data Augmentation
- On Mutual Information in Contrastive Learning for Visual Representations
- Debiased Contrastive Learning