Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
arXiv:2508.03860 · doi:10.1007/s10462-025-11454-w
Abstract
Large Language Models (LLMs) are trained on vast and diverse internet corpora that often include inaccurate or misleading content. Consequently, LLMs can generate misinformation, making robust fact-checking essential. This review systematically analyzes how LLM-generated content is evaluated for factual accuracy by exploring key challenges such as hallucinations, dataset limitations, and the reliability of evaluation metrics. The review emphasizes the need for strong fact-checking frameworks that integrate advanced prompting strategies, domain-specific fine-tuning, and retrieval-augmented generation (RAG) methods. It proposes five research questions that guide the analysis of the recent literature from 2020 to 2025, focusing on evaluation methods and mitigation techniques. Instruction tuning, multi-agent reasoning, and RAG frameworks for external knowledge access are also reviewed. The key findings demonstrate the limitations of current metrics, the importance of validated external evidence, and the improvement of factual consistency through domain-specific customization. The review underscores the importance of building more accurate, understandable, and context-aware fact-checking. These insights contribute to the advancement of research toward more trustworthy models.
References in corpus (10)
- Survey of Hallucination in Natural Language Generation
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- MDFEND: Multi-domain Fake News Detection
- Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
- Factuality Challenges in the Era of Large Language Models
- The Perils & Promises of Fact-checking with Large Language Models
- End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models
- Fact-checking information from large language models can decrease headline discernment
- EUvsDisinfo: A Dataset for Multilingual Detection of Pro-Kremlin Disinformation in News Articles
- MultiClaimNet: A Massively Multilingual Dataset of Fact-Checked Claim Clusters