Neural Semi-supervised Learning for Text Classification Under Large-Scale Pretraining
arXiv:2011.08626
Abstract
The goal of semi-supervised learning is to utilize the unlabeled, in-domain dataset U to improve models trained on the labeled dataset D. Under the context of large-scale language-model (LM) pretraining, how we can make the best use of U is poorly understood: is semi-supervised learning still beneficial with the presence of large-scale pretraining? should U be used for in-domain LM pretraining or pseudo-label generation? how should the pseudo-label based semi-supervised model be actually implemented? how different semi-supervised strategies affect performances regarding D of different sizes, U of different sizes, etc. In this paper, we conduct comprehensive studies on semi-supervised learning in the task of text classification under the context of large-scale LM pretraining. Our studies shed important lights on the behavior of semi-supervised learning methods: (1) with the presence of in-domain pretraining LM on U, open-domain LM pretraining is unnecessary; (2) both the in-domain pretraining strategy and the pseudo-label based strategy introduce significant performance boosts, with the former performing better with larger U, the latter performing better with smaller U, and the combination leading to the largest performance boost; (3) self-training (pretraining first on pseudo labels D' and then fine-tuning on D) yields better performances when D is small, while joint training on the combination of pseudo labels D' and the original dataset D yields better performances when D is large. Using semi-supervised learning strategies, we are able to achieve a performance of around 93.8% accuracy with only 50 training data points on the IMDB dataset, and a competitive performance of 96.6% with the full IMDB dataset. Our work marks an initial step in understanding the behavior of semi-supervised learning models under the context of large-scale pretraining.
References in corpus (27)
- Adam: A Method for Stochastic Optimization
- Distributed Representations of Words and Phrases and their Compositionality
- EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- Unsupervised Data Augmentation for Consistency Training
- Temporal Ensembling for Semi-Supervised Learning
- Interpolation Consistency Training for Semi-Supervised Learning
- AutoAugment: Learning Augmentation Policies from Data
- Finding Deceptive Opinion Spam by Any Stretch of the Imagination
- Semi-supervised Sequence Learning
- Big Self-Supervised Models are Strong Semi-Supervised Learners
- Rethinking Pre-training and Self-training
- Data Augmentation by Pairing Samples for Images Classification
- Billion-scale semi-supervised learning for image classification
- Learning to Self-Train for Semi-Supervised Few-Shot Classification
- Improved Noisy Student Training for Automatic Speech Recognition
- Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning
- Adversarial Training Methods for Semi-Supervised Text Classification
- Augmenting Data with Mixup for Sentence Classification: An Empirical Study
- Data Augmentation using Pre-trained Transformer Models
- XLDA: Cross-Lingual Data Augmentation for Natural Language Inference and Question Answering
- Soft Contextual Data Augmentation for Neural Machine Translation
- Low Resource Text Classification with ULMFit and Backtranslation
- Not Enough Data? Deep Learning to the Rescue!
- Description Based Text Classification with Reinforcement Learning
- Leveraging Just a Few Keywords for Fine-Grained Aspect Detection Through Weakly Supervised Co-Training