SCARF: Self-Supervised Contrastive Learning using Random Feature Corruption
arXiv:2106.15147
Abstract
Self-supervised contrastive representation learning has proved incredibly successful in the vision and natural language domains, enabling state-of-the-art performance with orders of magnitude less labeled data. However, such methods are domain-specific and little has been done to leverage this technique on real-world tabular datasets. We propose SCARF, a simple, widely-applicable technique for contrastive learning, where views are formed by corrupting a random subset of features. When applied to pre-train deep neural networks on the 69 real-world, tabular classification datasets from the OpenML-CC18 benchmark, SCARF not only improves classification accuracy in the fully-supervised setting but does so also in the presence of label noise and in the semi-supervised setting where only a fraction of the available training data is labeled. We show that SCARF complements existing strategies and outperforms alternatives like autoencoders. We conduct comprehensive ablations, detailing the importance of a range of factors.
ICLR 2022 Spotlight
References in corpus (16)
- Distilling the Knowledge in a Neural Network
- Natural Language Processing (almost) from Scratch
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- The Effectiveness of Data Augmentation in Image Classification using Deep Learning
- Cross-lingual Language Model Pretraining
- OpenML: networked science in machine learning
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Contrastive Multi-View Representation Learning on Graphs
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- BERTje: A Dutch BERT Model
- Population Based Augmentation: Efficient Learning of Augmentation Policy Schedules
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset Biases
- A Bayesian Data Augmentation Approach for Learning Deep Models
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- Viewmaker Networks: Learning Views for Unsupervised Representation Learning