Text Summarization with Pretrained Encoders
arXiv:1908.08345
Abstract
Bidirectional Encoder Representations from Transformers (BERT) represents the latest incarnation of pretrained language models which have recently advanced a wide range of natural language processing tasks. In this paper, we showcase how BERT can be usefully applied in text summarization and propose a general framework for both extractive and abstractive models. We introduce a novel document-level encoder based on BERT which is able to express the semantics of a document and obtain representations for its sentences. Our extractive model is built on top of this encoder by stacking several inter-sentence Transformer layers. For abstractive summarization, we propose a new fine-tuning schedule which adopts different optimizers for the encoder and the decoder as a means of alleviating the mismatch between the two (the former is pretrained while the latter is not). We also demonstrate that a two-staged fine-tuning approach can further boost the quality of the generated summaries. Experiments on three datasets show that our model achieves state-of-the-art results across the board in both extractive and abstractive settings. Our code is available at https://github.com/nlpyang/PreSumm
fix typos
References in corpus (4)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- SummaRuNNer: A Recurrent Neural Network based Sequence Model for Extractive Summarization of Documents
- Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
Cited by in corpus (47)
- Multilingual Denoising Pre-training for Neural Machine Translation
- Big Bird: Transformers for Longer Sequences
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
- Recursively Summarizing Books with Human Feedback
- Exploring the State of the Art in Legal QA Systems
- IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
- Robust Deep Reinforcement Learning for Extractive Legal Summarization
- BERT Fine-tuning For Arabic Text Summarization
- AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning
- Liputan6: A Large-scale Indonesian Dataset for Text Summarization
- Keyphrase Prediction With Pre-trained Language Model
- TED: A Pretrained Unsupervised Summarization Model with Theme Modeling and Denoising
- AdaptSum: Towards Low-Resource Domain Adaptation for Abstractive Summarization
- Text Summarization of Czech News Articles Using Named Entities
- Understanding Multi-Head Attention in Abstractive Summarization
- Dimsum @LaySumm 20: BART-based Approach for Scientific Document Summarization
- Enhancing Scientific Papers Summarization with Citation Graph
- Improving Faithfulness in Abstractive Summarization with Contrast Candidate Generation and Selection
- Product Title Generation for Conversational Systems using BERT
- Efficacy of BERT embeddings on predicting disaster from Twitter data
- ColdGANs: Taming Language GANs with Cautious Sampling Strategies
- Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries
- ProphetNet-X: Large-Scale Pre-training Models for English, Chinese, Multi-lingual, Dialog, and Code Generation
- Modelling Latent Skills for Multitask Language Generation
- Contextualized Rewriting for Text Summarization
- Unsupervised Pre-training for Natural Language Generation: A Literature Review
- An Experimental Evaluation of Transformer-based Language Models in the Biomedical Domain
- Augmented Abstractive Summarization With Document-LevelSemantic Graph
- SciSummPip: An Unsupervised Scientific Paper Summarization Pipeline
- Neural Entity Summarization with Joint Encoding and Weak Supervision
- Program Enhanced Fact Verification with Verbalization and Graph Attention Network
- Towards QoS-Aware and Resource-Efficient GPU Microservices Based on Spatial Multitasking GPUs In Datacenters
- Reference and Document Aware Semantic Evaluation Methods for Korean Language Summarization
- Lookup or Exploratory: What is Your Search Intent?
- CNewSum: A Large-scale Chinese News Summarization Dataset with Human-annotated Adequacy and Deducibility Level
- Fact-level Extractive Summarization with Hierarchical Graph Mask on BERT
- DeepTitle -- Leveraging BERT to generate Search Engine Optimized Headlines
- ScopeIt: Scoping Task Relevant Sentences in Documents
- VerbCL: A Dataset of Verbatim Quotes for Highlight Extraction in Case Law
- Explaining Documents' Relevance to Search Queries
- Improve Query Focused Abstractive Summarization by Incorporating Answer Relevance
- Topic Modeling and Progression of American Digital News Media During the Onset of the COVID-19 Pandemic
- Autoregressive Knowledge Distillation through Imitation Learning
- Exploratory Search with Sentence Embeddings
- CLAUSEREC: A Clause Recommendation Framework for AI-aided Contract Authoring
- Centrality Meets Centroid: A Graph-based Approach for Unsupervised Document Summarization