LexRank: Graph-based Lexical Centrality as Salience in Text Summarization
arXiv:1109.2128 · doi:10.1613/jair.1523
Abstract
We introduce a stochastic graph-based method for computing relative importance of textual units for Natural Language Processing. We test the technique on the problem of Text Summarization (TS). Extractive TS relies on the concept of sentence salience to identify the most important sentences in a document or set of documents. Salience is typically defined in terms of the presence of particular important words or in terms of similarity to a centroid pseudo-sentence. We consider a new approach, LexRank, for computing sentence importance based on the concept of eigenvector centrality in a graph representation of sentences. In this model, a connectivity matrix based on intra-sentence cosine similarity is used as the adjacency matrix of the graph representation of sentences. Our system, based on LexRank ranked in first place in more than one task in the recent DUC 2004 evaluation. In this paper we present a detailed analysis of our approach and apply it to a larger data set including data from earlier DUC evaluations. We discuss several methods to compute centrality using the similarity graph. The results show that degree-based methods (including LexRank) outperform both centroid-based methods and other systems participating in DUC in most of the cases. Furthermore, the LexRank with threshold method outperforms the other degree-based techniques including continuous LexRank. We also show that our approach is quite insensitive to the noise in the data that may result from an imperfect topical clustering of documents.
References in corpus (2)
Cited by in corpus (38)
- The Structure and Dynamics of Co-Citation Clusters: A Multiple-Perspective Co-Citation Analysis
- A Review of the Trends and Challenges in Adopting Natural Language Processing Methods for Education Feedback Analysis
- Information Retrieval: Recent Advances and Beyond
- Explainable Outfit Recommendation with Joint Outfit Matching and Comment Generation
- If I Hear You Correctly: Building and Evaluating Interview Chatbots with Active Listening Skills
- Using network science and text analytics to produce surveys in a scientific topic
- Abstractive Text Summarization: State of the Art, Challenges, and Improvements
- An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT
- sCAKE: Semantic Connectivity Aware Keyword Extraction
- Scientific document summarization via citation contextualization and scientific discourse
- Query-oriented text summarization based on hypergraph transversals
- Paragraph-based complex networks: application to document classification and authenticity verification
- Quantifying the informativeness for biomedical literature summarization: An itemset mining method
- Overview of BioASQ 2020: The eighth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
- Learning with fuzzy hypergraphs: a topical approach to query-oriented text summarization
- Extractive Summarization of Call Transcripts
- Investigating Entropy for Extractive Document Summarization
- Using Generic Summarization to Improve Music Information Retrieval Tasks
- A Survey on Event-based News Narrative Extraction
- Generative AI for Pull Request Descriptions: Adoption, Impact, and Developer Interventions
- MultiGBS: A multi-layer graph approach to biomedical summarization
- Summarization of Films and Documentaries Based on Subtitles and Scripts
- Summarization, Simplification, and Generation: The Case of Patents
- Dataset for Automatic Summarization of Russian News
- QBSUM: a Large-Scale Query-Based Document Summarization Dataset from Real-world Applications
- A System for Interleaving Discussion and Summarization in Online Collaboration
- Keyphrase Generation: A Multi-Aspect Survey
- An entity-guided text summarization framework with relational heterogeneous graph neural network
- Exploring Optimal Granularity for Extractive Summarization of Unstructured Health Records: Analysis of the Largest Multi-Institutional Archive of Health Records in Japan
- Generalized minimum dominating set and application in automatic text summarization
- Advancements in Natural Language Processing for Automatic Text Summarization
- Semantic Similarity Measure of Natural Language Text through Machine Learning and a Keyword-Aware Cross-Encoder-Ranking Summarizer -- A Case Study Using UCGIS GIS&T Body of Knowledge
- Textual analysis of End User License Agreement for red-flagging potentially malicious software
- A Hybrid Architecture with Efficient Fine Tuning for Abstractive Patent Document Summarization
- Using Query Expansion in Manifold Ranking for Query-Oriented Multi-Document Summarization
- An Information-theoretic Approach to Machine-oriented Music Summarization
- Development of an Extractive Title Generation System Using Titles of Papers of Top Conferences for Intermediate English Students
- A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs