Temporal Ensembling for Semi-Supervised Learning
arXiv:1610.02242
Abstract
In this paper, we present a simple and efficient method for training deep neural networks in a semi-supervised setting where only a small portion of training data is labeled. We introduce self-ensembling, where we form a consensus prediction of the unknown labels using the outputs of the network-in-training on different epochs, and most importantly, under different regularization and input augmentation conditions. This ensemble prediction can be expected to be a better predictor for the unknown labels than the output of the network at the most recent training epoch, and can thus be used as a target for training. Using our method, we set new records for two standard semi-supervised learning benchmarks, reducing the (non-augmented) classification error rate from 18.44% to 7.05% in SVHN with 500 labels and from 18.63% to 16.55% in CIFAR-10 with 4000 labels, and further to 5.12% and 12.16% by enabling the standard augmentations. We additionally obtain a clear improvement in CIFAR-100 classification accuracy by using random images from the Tiny Images dataset as unlabeled extra inputs during training. Finally, we demonstrate good tolerance to incorrect labels.
Final ICLR 2017 version. Includes new results for CIFAR-100 with additional unlabeled data from Tiny Images dataset
References in corpus (6)
- Distilling the Knowledge in a Neural Network
- Striving for Simplicity: The All Convolutional Net
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- Training Convolutional Networks with Noisy Labels
- Fractional Max-Pooling
- Making Deep Neural Networks Robust to Label Noise: a Loss Correction Approach
Cited by in corpus (51)
- Bootstrap your own latent: A new approach to self-supervised Learning
- Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
- Rethinking the Value of Labels for Improving Class-Imbalanced Learning
- Image Augmentations for GAN Training
- FDA: Fourier Domain Adaptation for Semantic Segmentation
- SELF: Learning to Filter Noisy Labels with Self-Ensembling
- OSLNet: Deep Small-Sample Classification with an Orthogonal Softmax Layer
- Consistency Regularization for Generative Adversarial Networks
- Class-Imbalanced Semi-Supervised Learning
- Deep Two-path Semi-supervised Learning for Fake News Detection
- Structured Generative Adversarial Networks
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification
- Learning Kernel for Conditional Moment-Matching Discrepancy-based Image Classification
- Contrastive and Generative Graph Convolutional Networks for Graph-based Semi-Supervised Learning
- Semi-Supervised Brain Lesion Segmentation with an Adapted Mean Teacher Model
- FeatMatch: Feature-Based Augmentation for Semi-Supervised Learning
- Semi-Supervised Learning with Normalizing Flows
- Learning from Label Proportions with Consistency Regularization
- Extreme Consistency: Overcoming Annotation Scarcity and Domain Shifts
- AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations
- 3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning
- Semi-supervised Learning using Adversarial Training with Good and Bad Samples
- Semi-Supervised Learning for Fetal Brain MRI Quality Assessment with ROI consistency
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- Cell Type Identification from Single-Cell Transcriptomic Data via Semi-supervised Learning
- Semi-Supervised Learning for In-Game Expert-Level Music-to-Dance Translation
- Non-technical Loss Detection with Statistical Profile Images Based on Semi-supervised Learning
- Label Prediction Framework for Semi-Supervised Cross-Modal Retrieval
- A Comprehensive Approach to Unsupervised Embedding Learning based on AND Algorithm
- SSAH: Semi-supervised Adversarial Deep Hashing with Self-paced Hard Sample Generation
- Mixup-breakdown: a consistency training method for improving generalization of speech separation models
- Consistency Regularization with Generative Adversarial Networks for Semi-Supervised Learning
- Temporal Self-Ensembling Teacher for Semi-Supervised Object Detection
- Target Consistency for Domain Adaptation: when Robustness meets Transferability
- Optimally Combining Classifiers for Semi-Supervised Learning
- Metric learning by Similarity Network for Deep Semi-Supervised Learning
- Learning to Detect Important People in Unlabelled Images for Semi-supervised Important People Detection
- Semi-Supervised Learning with IPM-based GANs: an Empirical Study
- Weakly Supervised Person Re-Identification
- Learning Temporal Action Proposals With Fewer Labels
- Attribute-Induced Bias Eliminating for Transductive Zero-Shot Learning
- Active Crowd Counting with Limited Supervision
- Improving Object Detection with Selective Self-supervised Self-training
- Pairwise Teacher-Student Network for Semi-Supervised Hashing
- Long Short-Term Sample Distillation
- Domain Constraint Approximation based Semi Supervision
- Dual-Teacher: Integrating Intra-domain and Inter-domain Teachers for Annotation-efficient Cardiac Segmentation
- Understanding Classifier Mistakes with Generative Models
- A Self-Supervised Bootstrap Method for Single-Image 3D Face Reconstruction
- CFEA: Collaborative Feature Ensembling Adaptation for Domain Adaptation in Unsupervised Optic Disc and Cup Segmentation
- AMC-Loss: Angular Margin Contrastive Loss for Improved Explainability in Image Classification