News Summarization and Evaluation in the Era of GPT-3
arXiv:2209.12356
Abstract
The recent success of prompting large language models like GPT-3 has led to a paradigm shift in NLP research. In this paper, we study its impact on text summarization, focusing on the classic benchmark domain of news summarization. First, we investigate how GPT-3 compares against fine-tuned models trained on large summarization datasets. We show that not only do humans overwhelmingly prefer GPT-3 summaries, prompted using only a task description, but these also do not suffer from common dataset-specific issues such as poor factuality. Next, we study what this means for evaluation, particularly the role of gold standard test sets. Our experiments show that both reference-based and reference-free automatic metrics cannot reliably evaluate GPT-3 summaries. Finally, we evaluate models on a setting beyond generic summarization, specifically keyword-based summarization, and show how dominant fine-tuning approaches compare to prompting. To support further research, we release: (a) a corpus of 10K generated summaries from fine-tuned and prompt-based models across 4 standard summarization benchmarks, (b) 1K human preference judgments comparing different systems for generic- and keyword-based summarization.
All data shared at: https://tagoyal.github.io/zeroshot-news-annotations.html
Cited by in corpus (11)
- ChatGPT: Jack of all trades, master of none
- The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors
- Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
- LEVA: Using Large Language Models to Enhance Visual Analytics
- Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation
- A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital-Course Summarization
- Multi-Dimensional Evaluation of Text Summarization with In-Context Learning
- Translating Legalese: Enhancing Public Understanding of Court Opinions with Legal Summarizers
- GPT Struct Me: Probing GPT Models on Narrative Entity Extraction
- Precise Length Control in Large Language Models
- The Viability of Crowdsourcing for RAG Evaluation