EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks
arXiv:1901.11196
Abstract
We present EDA: easy data augmentation techniques for boosting performance on text classification tasks. EDA consists of four simple but powerful operations: synonym replacement, random insertion, random swap, and random deletion. On five text classification tasks, we show that EDA improves performance for both convolutional and recurrent neural networks. EDA demonstrates particularly strong results for smaller datasets; on average, across five datasets, training with EDA while using only 50% of the available training set achieved the same accuracy as normal training with all available data. We also performed extensive ablation studies and suggest parameters for practical use.
EMNLP-IJCNLP 2019 short paper
References in corpus (5)
Cited by in corpus (23)
- Data Augmentation using Pre-trained Transformer Models
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- MixKD: Towards Efficient Distillation of Large-scale Language Models
- DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
- DeepSentiPers: Novel Deep Learning Models Trained Over Proposed Augmented Persian Sentiment Corpus
- Effects of Word-frequency based Pre- and Post- Processings for Audio Captioning
- Semantic Equivalent Adversarial Data Augmentation for Visual Question Answering
- Cost-Sensitive BERT for Generalisable Sentence Classification with Imbalanced Data
- Enhanced Offensive Language Detection Through Data Augmentation
- Composed Variational Natural Language Generation for Few-shot Intents
- Privacy-preserving Collaborative Learning with Automatic Transformation Search
- Joint System-Wise Optimization for Pipeline Goal-Oriented Dialog System
- Identifying Introductions in Podcast Episodes from Automatically Generated Transcripts
- An Approach to Improve Robustness of NLP Systems against ASR Errors
- Sexism Identification in Tweets and Gabs using Deep Neural Networks
- The effects of data size on Automated Essay Scoring engines
- Linguistic Knowledge in Data Augmentation for Natural Language Processing: An Example on Chinese Question Matching
- CERM: Context-aware Literature-based Discovery via Sentiment Analysis
- Pseudo Siamese Network for Few-shot Intent Generation
- SDA: Improving Text Generation with Self Data Augmentation
- A Transformer Based Pitch Sequence Autoencoder with MIDI Augmentation
- Automatically Select Emotion for Response via Personality-affected Emotion Transition
- A Diversity-Enhanced and Constraints-Relaxed Augmentation for Low-Resource Classification