Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
arXiv:1907.12461 · doi:10.1162/tacl_a_00313
Abstract
Unsupervised pre-training of large neural models has recently revolutionized Natural Language Processing. By warm-starting from the publicly released checkpoints, NLP practitioners have pushed the state-of-the-art on multiple benchmarks while saving significant amounts of compute time. So far the focus has been mainly on the Natural Language Understanding tasks. In this paper, we demonstrate the efficacy of pre-trained checkpoints for Sequence Generation. We developed a Transformer-based sequence-to-sequence model that is compatible with publicly available pre-trained BERT, GPT-2 and RoBERTa checkpoints and conducted an extensive empirical study on the utility of initializing our model, both encoder and decoder, with these checkpoints. Our models result in new state-of-the-art results on Machine Translation, Text Summarization, Sentence Splitting, and Sentence Fusion.
To be published in Transactions of the Association for Computational Linguistics (TACL)
References in corpus (6)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- KERMIT: Generative Insertion-Based Modeling for Sequences
- Sample Efficient Text Summarization Using a Single Pre-Trained Transformer
- DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion
- To Tune or Not To Tune? How About the Best of Both Worlds?
Cited by in corpus (34)
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
- Big Bird: Transformers for Longer Sequences
- Exploring Chemical Space using Natural Language Processing Methodologies for Drug Discovery
- Improving Chest X-Ray Report Generation by Leveraging Warm Starting
- Text Summarization with Pretrained Encoders
- IndicBART: A Pre-trained Model for Indic Natural Language Generation
- Discharge Summary Hospital Course Summarisation of In Patient Electronic Health Record Text with Clinical Concept Guided Deep Pre-Trained Transformer Models
- Sticking to the Facts: Confident Decoding for Faithful Data-to-Text Generation
- Learning to summarize from human feedback
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- ETC: Encoding Long and Structured Inputs in Transformers
- ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation
- E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
- Leveraging ParsBERT and Pretrained mT5 for Persian Abstractive Text Summarization
- GO FIGURE: A Meta Evaluation of Factuality in Summarization
- Detecting Hallucinated Content in Conditional Neural Sequence Generation
- NaturalProofs: Mathematical Theorem Proving in Natural Language
- Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
- Leveraging Graph to Improve Abstractive Multi-Document Summarization
- GLGE: A New General Language Generation Evaluation Benchmark
- Controllable Text Simplification with Explicit Paraphrasing
- Encoder-Decoder Models Can Benefit from Pre-trained Masked Language Models in Grammatical Error Correction
- Generating Representative Headlines for News Stories
- Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation
- VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding
- Neural CRF Model for Sentence Alignment in Text Simplification
- Code-switching pre-training for neural machine translation
- Multi-stage Pretraining for Abstractive Summarization
- Unsupervised Pre-training for Natural Language Generation: A Literature Review
- ReadTwice: Reading Very Large Documents with Memories
- BERT got a Date: Introducing Transformers to Temporal Tagging
- Multilingual AMR-to-Text Generation
- Semantically Driven Sentence Fusion: Modeling and Evaluation
- Sentence-Permuted Paragraph Generation