Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries
arXiv:2010.07074
Abstract
The difficulty of generating coherent long texts lies in the fact that existing models overwhelmingly focus on predicting local words, and cannot make high level plans on what to generate or capture the high-level discourse dependencies between chunks of texts. Inspired by human writing processes, where a list of bullet points or a catalog is first outlined, and then each bullet point is expanded to form the whole article, we propose {\it SOE}, a pipelined system that involves of summarizing, outlining and elaborating for long text generation: the model first outlines the summaries for different segments of long texts, and then elaborates on each bullet point to generate the corresponding segment. To avoid the labor-intensive process of summary soliciting, we propose the {\it reconstruction} strategy, which extracts segment summaries in an unsupervised manner by selecting its most informative part to reconstruct the segment. The proposed generation system comes with the following merits: (1) the summary provides high-level guidance for text generation and avoids the local minimum of individual word predictions; (2) the high-level discourse dependencies are captured in the conditional dependencies between summaries and are preserved during the summary expansion process and (3) additionally, we are able to consider significantly more contexts by representing contexts as concise summaries. Extensive experiments demonstrate that SOE produces long texts with significantly better quality, along with faster convergence speed.
To appear at COLING 2022
References in corpus (16)
- Language Models are Few-Shot Learners
- Generating Long Sequences with Sparse Transformers
- Pointer Sentinel Mixture Models
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- A Hierarchical Neural Autoencoder for Paragraphs and Documents
- Data Noising as Smoothing in Neural Network Language Models
- Long Text Generation via Adversarial Training with Leaked Information
- Generating Wikipedia by Summarizing Long Sequences
- Adversarial Evaluation of Dialogue Models
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning
- Challenges in Data-to-Document Generation
- Extractive Summarization as Text Matching
- Step-by-Step: Separating Planning from Realization in Neural Data-to-Text Generation
- Jointly Measuring Diversity and Quality in Text Generation Models
- Towards Generating Long and Coherent Text with Multi-Level Latent Variable Models