Theoretical Analysis of Self-Training with Deep Networks on Unlabeled Data
arXiv:2010.03622
Abstract
Self-training algorithms, which train a model to fit pseudolabels predicted by another previously-learned model, have been very successful for learning with unlabeled data using neural networks. However, the current theoretical understanding of self-training only applies to linear models. This work provides a unified theoretical analysis of self-training with deep networks for semi-supervised learning, unsupervised domain adaptation, and unsupervised learning. At the core of our analysis is a simple but realistic "expansion" assumption, which states that a low probability subset of the data must expand to a neighborhood with large probability relative to the subset. We also assume that neighborhoods of examples in different classes have minimal overlap. We prove that under these assumptions, the minimizers of population objectives based on self-training and input-consistency regularization will achieve high accuracy with respect to ground-truth labels. By using off-the-shelf generalization bounds, we immediately convert this result to sample complexity guarantees for neural nets that are polynomial in the margin and Lipschitzness. Our results help explain the empirical successes of recently proposed self-training algorithms which use input consistency regularization.
Published at ICLR 2021
References in corpus (10)
- Semi-Supervised Classification with Graph Convolutional Networks
- Bootstrap your own latent: A new approach to self-supervised Learning
- Deep Domain Confusion: Maximizing for Domain Invariance
- Semi-Supervised Learning with Deep Generative Models
- Billion-scale semi-supervised learning for image classification
- Bridging Theory and Algorithm for Domain Adaptation
- Generalization error bounds in semi-supervised classification under the cluster assumption
- Predicting What You Already Know Helps: Provable Self-Supervised Learning
- Deterministic PAC-Bayesian generalization bounds for deep networks via generalizing noise-resilience
- Contrastive estimation reveals topic posterior information to linear models
Cited by in corpus (23)
- On the Opportunities and Risks of Foundation Models
- Self-Training: A Survey
- Cycle Self-Training for Domain Adaptation
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive Loss
- SemiFL: Semi-Supervised Federated Learning for Unlabeled Clients with Alternate Training
- Contrastive Self-supervised Neural Architecture Search
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt Tuning
- Towards the Generalization of Contrastive Self-Supervised Learning
- USCORE: An Effective Approach to Fully Unsupervised Evaluation Metrics for Machine Translation
- Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
- Understanding Negative Samples in Instance Discriminative Self-supervised Representation Learning
- A Theory of Label Propagation for Subpopulation Shift
- Disambiguation of weak supervision with exponential convergence rates
- Can Pretext-Based Self-Supervised Learning Be Boosted by Downstream Data? A Theoretical Analysis
- Streaming Self-Training via Domain-Agnostic Unlabeled Images
- Incompatibility Clustering as a Defense Against Backdoor Poisoning Attacks
- Online Continual Adaptation with Active Self-Training
- Multimodal Knowledge Expansion
- Task-adaptive Pre-training and Self-training are Complementary for Natural Language Understanding
- CoDiM: Learning with Noisy Labels via Contrastive Semi-Supervised Learning
- Combining Diverse Feature Priors
- Pretext Tasks selection for multitask self-supervised speech representation learning
- Self-Tuning for Data-Efficient Deep Learning