Self-Adaptive Training: Bridging Supervised and Self-Supervised Learning
arXiv:2101.08732
Abstract
We propose self-adaptive training -- a unified training algorithm that dynamically calibrates and enhances training processes by model predictions without incurring an extra computational cost -- to advance both supervised and self-supervised learning of deep neural networks. We analyze the training dynamics of deep networks on training data that are corrupted by, e.g., random noise and adversarial examples. Our analysis shows that model predictions are able to magnify useful underlying information in data and this phenomenon occurs broadly even in the absence of any label information, highlighting that model predictions could substantially benefit the training processes: self-adaptive training improves the generalization of deep networks under noise and enhances the self-supervised representation learning. The analysis also sheds light on understanding deep learning, e.g., a potential explanation of the recently-discovered double-descent phenomenon in empirical risk minimization and the collapsing issue of the state-of-the-art self-supervised learning algorithms. Experiments on the CIFAR, STL, and ImageNet datasets verify the effectiveness of our approach in three applications: classification with label noise, selective classification, and linear evaluation. To facilitate future research, the code has been made publicly available at https://github.com/LayneH/self-adaptive-training.
Accepted at T-PAMI. Journal version of arXiv:2002.10319 [cs.LG] (NeurIPS2020). 22 pages, 15 figures, 13 tables
References in corpus (14)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Theoretically Principled Trade-off between Robustness and Accuracy
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- Prototypical Contrastive Learning of Unsupervised Representations
- A Closer Look at Memorization in Deep Networks
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- Uniform convergence may be unable to explain generalization in deep learning
- Combating Label Noise in Deep Learning Using Abstention
- Unsupervised Deep Learning by Neighbourhood Discovery
- SelectiveNet: A Deep Neural Network with an Integrated Reject Option
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network