A Survey on Data Augmentation for Text Classification
arXiv:2107.03158 · doi:10.1145/3544558
Abstract
Data augmentation, the artificial creation of training data for machine learning by transformations, is a widely studied research field across machine learning disciplines. While it is useful for increasing a model's generalization capabilities, it can also address many other challenges and problems, from overcoming a limited amount of training data, to regularizing the objective, to limiting the amount data used to protect privacy. Based on a precise description of the goals and applications of data augmentation and a taxonomy for existing works, this survey is concerned with data augmentation methods for textual classification and aims to provide a concise and comprehensive overview for researchers and practitioners. Derived from the taxonomy, we divide more than 100 methods into 12 different groupings and give state-of-the-art references expounding which methods are highly promising by relating them to each other. Finally, research perspectives that may constitute a building block for future work are provided.
44 pages, 5 figures, 9 tables
References in corpus (20)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Language Models are Few-Shot Learners
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- Assessing BERT's Syntactic Abilities
- CLEAR: Contrastive Learning for Sentence Representation
- Data Noising as Smoothing in Neural Network Language Models
- Data Augmentation in Natural Language Processing: A Novel Text Generation Approach for Long and Short Text Classifiers
- Augmenting Data with Mixup for Sentence Classification: An Empirical Study
- What do you learn from context? Probing for sentence structure in contextualized word representations
- A Simple but Tough-to-Beat Data Augmentation Approach for Natural Language Understanding and Generation
- Adversarial Training for Large Neural Language Models
- Learning Data Manipulation for Augmentation and Weighting
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional Generation
- Soft Contextual Data Augmentation for Neural Machine Translation
- Neural Data Augmentation via Example Extrapolation
- CoDA: Contrast-enhanced and Diversity-promoting Data Augmentation for Natural Language Understanding
- AugLy: Data Augmentations for Robustness
- NL-Augmenter: A Framework for Task-Sensitive Natural Language Augmentation
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text Classification
- Empirical Study of Text Augmentation on Social Media Text in Vietnamese
Cited by in corpus (10)
- Data Augmentation Approaches in Natural Language Processing: A Survey
- Data Augmentation in Natural Language Processing: A Novel Text Generation Approach for Long and Short Text Classifiers
- Significantly improving zero-shot X-ray pathology classification via fine-tuning pre-trained image-text encoders
- Explainable AI: XAI-Guided Context-Aware Data Augmentation
- Gender Bias Detection in Court Decisions: A Brazilian Case Study
- Evaluating the Impact of Data Augmentation on Predictive Model Performance
- AI Simulation by Digital Twins: Systematic Survey, Reference Framework, and Mapping to a Standardized Architecture
- Reducing and Exploiting Data Augmentation Noise through Meta Reweighting Contrastive Learning for Text Classification
- Revisiting data augmentation for subspace clustering
- EmoGRACE: Aspect-based emotion analysis for social media data