natural language processing

Query-Focused Event Summarization: A Dataset and Benchmark

arXiv:2607.11166

summary

The paper defines a new Query-Focused Event Summarization (QFES) task, releases a large QFESum dataset covering multiple thematic events, and proposes a two‑stage framework (adaptive retrieval and hierarchical clustering summarization) that outperforms existing baselines.

Abstract

A thematic corpus is a collection of semantically coherent documents that collectively describe different aspects of a shared thematic event. Such a corpus typically contains hundreds or even thousands of documents. While users' interests in a thematic event often span multiple dimensions, Query-Focused Summarization (QFS) aims to generate summaries tailored to users' queries. However, existing QFS datasets lack event-oriented summarization, and most QFS methods struggle with large-scale corpora. To address these challenges, we propose the Query-Focused Event Summarization (QFES) task and construct the QFESum dataset, which contains 8 thematic events, 16,684 documents, and 104 queries. Furthermore, we introduce a two-stage QFES framework consisting of Query-Focused Retrieval with Adaptive Thresholding (RAT) and Query-Focused Summarization based on Hierarchical Clustering (SHC). Experimental results on QFESum show that RAT and SHC consistently outperform the baselines, demonstrating their effectiveness for QFES. The dataset and code are publicly available at https://github.com/sarcasm-hcy02/QFES-QFESum.

22 pages, 9 figures, and 13 tables. Dataset and code are available at https://github.com/sarcasm-hcy02/QFES-QFESum

Topics & keywords

#query-focused summarization#event summarization#large-scale corpora#information retrieval#hierarchical clusteringQFESum datasetadaptive thresholdinghierarchical clustering summarizationquery-focused retrievalthematic event