Data Augmentation using Pre-trained Transformer Models
arXiv:2003.02245
Abstract
Language model based pre-trained models such as BERT have provided significant gains across different NLP tasks. In this paper, we study different types of transformer based pre-trained models such as auto-regressive models (GPT-2), auto-encoder models (BERT), and seq2seq models (BART) for conditional data augmentation. We show that prepending the class labels to text sequences provides a simple yet effective way to condition the pre-trained models for data augmentation. Additionally, on three classification benchmarks, pre-trained Seq2Seq model outperforms other data augmentation methods in a low-resource setting. Further, we explore how different pre-trained model based data augmentation differs in-terms of data diversity, and how well such methods preserve the class-label information.
In Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems @ AACL 2020; Code: https://github.com/varinf/TransformersDataAugmentation
References in corpus (10)
- Adam: A Method for Stochastic Optimization
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Unsupervised Data Augmentation for Consistency Training
- The Curious Case of Neural Text Degeneration
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces
- EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks
- Understanding Back-Translation at Scale
- Learning Data Manipulation for Augmentation and Weighting
- Low Resource Text Classification with ULMFit and Backtranslation
Cited by in corpus (29)
- Pre-trained Models for Natural Language Processing: A Survey
- A Survey on Data Augmentation for Text Classification
- Data Augmentation Approaches in Natural Language Processing: A Survey
- Big Bird: Transformers for Longer Sequences
- Data Augmentation in Natural Language Processing: A Novel Text Generation Approach for Long and Short Text Classifiers
- Generative Data Augmentation for Commonsense Reasoning
- Improving Short Text Classification With Augmented Data Using GPT-3
- Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions
- Towards Zero-Label Language Learning
- On Data Augmentation for Extreme Multi-label Classification
- EIGEN: Event Influence GENeration using Pre-trained Language Models
- Simple is Better! Lightweight Data Augmentation for Low Resource Slot Filling and Intent Classification
- Unsupervised Neural Machine Translation with Generative Language Models Only
- Neural Semi-supervised Learning for Text Classification Under Large-Scale Pretraining
- DAGA: Data Augmentation with a Generation Approach for Low-resource Tagging Tasks
- Fine-tuning of Pre-trained Transformers for Hate, Offensive, and Profane Content Detection in English and Marathi
- Say What? Collaborative Pop Lyric Generation Using Multitask Transfer Learning
- AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data
- BET: A Backtranslation Approach for Easy Data Augmentation in Transformer-based Paraphrase Identification Context
- Text Data Augmentation: Towards better detection of spear-phishing emails
- Investigation on Data Adaptation Techniques for Neural Named Entity Recognition
- Automated Annotation of Scientific Texts for ML-based Keyphrase Extraction and Validation
- Context-gloss Augmentation for Improving Word Sense Disambiguation
- UU-Tax at SemEval-2022 Task 3: Improving the generalizability of language models for taxonomy classification through data augmentation
- Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous Space
- Generating Synthetic Data for Task-Oriented Semantic Parsing with Hierarchical Representations
- Gradient Imitation Reinforcement Learning for Low Resource Relation Extraction
- Simulated Chats for Building Dialog Systems: Learning to Generate Conversations from Instructions
- Data Augmentation for Spoken Language Understanding via Pretrained Language Models