Assessing The Factual Accuracy of Generated Text
arXiv:1905.13322 · doi:10.1145/3292500.3330955
Abstract
We propose a model-based metric to estimate the factual accuracy of generated text that is complementary to typical scoring schemes like ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and BLEU (Bilingual Evaluation Understudy). We introduce and release a new large-scale dataset based on Wikipedia and Wikidata to train relation classifiers and end-to-end fact extraction models. The end-to-end models are shown to be able to extract complete sets of facts from datasets with full pages of text. We then analyse multiple models that estimate factual accuracy on a Wikipedia text summarization task, and show their efficacy compared to ROUGE and other model-free variants by conducting a human evaluation study.
References in corpus (3)
Cited by in corpus (35)
- Survey of Hallucination in Natural Language Generation
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- An Empirical Survey on Long Document Summarization: Datasets, Models and Metrics
- Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions
- The Factual Inconsistency Problem in Abstractive Text Summarization: A Survey
- Asking and Answering Questions to Evaluate the Factual Consistency of Summaries
- Constrained Abstractive Summarization: Preserving Factual Consistency with Constrained Generation
- GO FIGURE: A Meta Evaluation of Factuality in Summarization
- Reasoning Over Semantic-Level Graph for Fact Checking
- CoCon: A Self-Supervised Approach for Controlled Text Generation
- Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation
- Enhancing Factual Consistency of Abstractive Summarization
- Challenges in Domain-Specific Abstractive Summarization and How to Overcome them
- Factual Error Correction for Abstractive Summarization Models
- A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital-Course Summarization
- Logical Natural Language Generation from Open-Domain Tables
- LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
- CO2Sum:Contrastive Learning for Factual-Consistent Abstractive Summarization
- Reducing Quantity Hallucinations in Abstractive Summarization
- TODSum: Task-Oriented Dialogue Summarization with State Tracking
- Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports
- WebRED: Effective Pretraining And Finetuning For Relation Extraction On The Web
- Resource for Error Analysis in Text Simplification: New Taxonomy and Test Collection
- Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric View
- Multi-Fact Correction in Abstractive Text Summarization
- Neural Deepfake Detection with Factual Structure of Text
- Is In-hospital Meta-information Useful for Abstractive Discharge Summary Generation?
- Information Retrieval in the Age of Generative AI: The RGB Model
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering
- Program Enhanced Fact Verification with Verbalization and Graph Attention Network
- Toward Improving Coherence and Diversity of Slogan Generation
- Exploring Decomposition for Table-based Fact Verification
- Few Shot Learning for Information Verification
- Factual Consistency Evaluation for Text Summarization via Counterfactual Estimation
- LogicalFactChecker: Leveraging Logical Operations for Fact Checking with Graph Module Network