Abstractive Text Summarization: State of the Art, Challenges, and Improvements
arXiv:2409.02413 · doi:10.1016/j.neucom.2024.128255
Abstract
Specifically focusing on the landscape of abstractive text summarization, as opposed to extractive techniques, this survey presents a comprehensive overview, delving into state-of-the-art techniques, prevailing challenges, and prospective research directions. We categorize the techniques into traditional sequence-to-sequence models, pre-trained large language models, reinforcement learning, hierarchical methods, and multi-modal summarization. Unlike prior works that did not examine complexities, scalability and comparisons of techniques in detail, this review takes a comprehensive approach encompassing state-of-the-art methods, challenges, solutions, comparisons, limitations and charts out future improvements - providing researchers an extensive overview to advance abstractive summarization research. We provide vital comparison tables across techniques categorized - offering insights into model complexity, scalability and appropriate applications. The paper highlights challenges such as inadequate meaning representation, factual consistency, controllable text summarization, cross-lingual summarization, and evaluation metrics, among others. Solutions leveraging knowledge incorporation and other innovative strategies are proposed to address these challenges. The paper concludes by highlighting emerging research areas like factual inconsistency, domain-specific, cross-lingual, multilingual, and long-document summarization, as well as handling noisy data. Our objective is to provide researchers and practitioners with a structured overview of the domain, enabling them to better understand the current landscape and identify potential areas for further research and improvement.
9 Tables, 7 Figures
References in corpus (24)
- Efficient Estimation of Word Representations in Vector Space
- Neural Machine Translation by Jointly Learning to Align and Translate
- Sequence to Sequence Learning with Neural Networks
- Survey of Hallucination in Natural Language Generation
- Towards A Rigorous Science of Interpretable Machine Learning
- A Survey on Knowledge Graphs: Representation, Acquisition and Applications
- BERTScore: Evaluating Text Generation with BERT
- LexRank: Graph-based Lexical Centrality as Salience in Text Summarization
- A Deep Reinforced Model for Abstractive Summarization
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- A Survey Of Cross-lingual Word Embedding Models
- How2: A Large-scale Dataset for Multimodal Language Understanding
- Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
- Automatic Text Summarization of COVID-19 Medical Research Articles using BERT and GPT-2
- Improving deep neural network generalization and robustness to background bias via layer-wise relevance propagation optimization
- Actor-Critic based Training Framework for Abstractive Summarization
- Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
- Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
- Multi-modal Summarization for Video-containing Documents
- Multi-Document Scientific Summarization from a Knowledge Graph-Centric View
- Personalized Abstractive Summarization by Tri-agent Generation Pipeline
- A Survey on Multi-modal Summarization
- Integrating Summarization and Retrieval for Enhanced Personalization via Large Language Models