Boosting Summarization with Normalizing Flows and Aggressive Training
arXiv:2311.00588 · doi:10.18653/v1/2023.emnlp-main.165
Abstract
This paper presents FlowSUM, a normalizing flows-based variational encoder-decoder framework for Transformer-based summarization. Our approach tackles two primary challenges in variational summarization: insufficient semantic information in latent representations and posterior collapse during training. To address these challenges, we employ normalizing flows to enable flexible latent posterior modeling, and we propose a controlled alternate aggressive training (CAAT) strategy with an improved gate mechanism. Experimental results show that FlowSUM significantly enhances the quality of generated summaries and unleashes the potential for knowledge distillation with minimal impact on inference time. Furthermore, we investigate the issue of posterior collapse in normalizing flows and analyze how the summary quality is affected by the training strategy, gate initialization, and the type and number of normalizing flows used, offering valuable insights for future research.
References in corpus (15)
- Distilling the Knowledge in a Neural Network
- BERTScore: Evaluating Text Generation with BERT
- NICE: Non-linear Independent Components Estimation
- A Deep Reinforced Model for Abstractive Summarization
- The Curious Case of Neural Text Degeneration
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- InfoVAE: Information Maximizing Variational Autoencoders
- Lagging Inference Networks and Posterior Collapse in Variational Autoencoders
- MeanSum: A Neural Model for Unsupervised Multi-document Abstractive Summarization
- Pre-trained Summarization Distillation
- Discrete Flows: Invertible Generative Models of Discrete Data
- Improving the Gating Mechanism of Recurrent Neural Networks
- Topic-Guided Abstractive Text Summarization: a Joint Learning Approach
- Diverse Text Generation via Variational Encoder-Decoder Models with Gaussian Process Priors