Factuality Challenges in the Era of Large Language Models
arXiv:2310.05189 · doi:10.1038/s42256-024-00881-z
Abstract
The emergence of tools based on Large Language Models (LLMs), such as OpenAI's ChatGPT, Microsoft's Bing Chat, and Google's Bard, has garnered immense public attention. These incredibly useful, natural-sounding tools mark significant advances in natural language generation, yet they exhibit a propensity to generate false, erroneous, or misleading content -- commonly referred to as "hallucinations." Moreover, LLMs can be exploited for malicious applications, such as generating false but credible-sounding content and profiles at scale. This poses a significant challenge to society in terms of the potential deception of users and the increasing dissemination of inaccurate information. In light of these risks, we explore the kinds of technological innovations, regulatory reforms, and AI literacy initiatives needed from fact-checkers, news organizations, and the broader research and policy communities. By identifying the risks, the imminent threats, and some viable solutions, we seek to shed light on navigating various aspects of veracity in the era of generative AI.
Our article offers a comprehensive examination of the challenges and risks associated with Large Language Models (LLMs), focusing on their potential impact on the veracity of information in today's digital landscape
References in corpus (27)
- Survey of Hallucination in Natural Language Generation
- LLaMA: Open and Efficient Foundation Language Models
- ChatGPT: Jack of all trades, master of none
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
- Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
- Carbon Emissions and Large Neural Network Training
- Exposure to Social Engagement Metrics Increases Vulnerability to Misinformation
- The Perils & Promises of Fact-checking with Large Language Models
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Anatomy of an AI-powered malicious social botnet
- Factuality Enhanced Language Models for Open-Ended Text Generation
- Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
- FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
- On the Reliability of Watermarks for Large Language Models
- Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models
- Characteristics and prevalence of fake social media profiles with AI-generated faces
- M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
- FELM: Benchmarking Factuality Evaluation of Large Language Models
- Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video
- Evaluating Hallucinations in Chinese Large Language Models
- R-Tuning: Instructing Large Language Models to Say `I Don't Know'
- Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language Models
- SmartBook: AI-Assisted Situation Report Generation for Intelligence Analysts
- FLAME: Factuality-Aware Alignment for Large Language Models
- EVEDIT: Event-based Knowledge Editing with Deductive Editing Boundaries
Cited by in corpus (11)
- Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
- Information Suppression in Large Language Models: Auditing, Quantifying, and Characterizing Censorship in DeepSeek
- A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language Models
- Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
- "There Has To Be a Lot That We're Missing": Moderating AI-Generated Content on Reddit
- Investigating the heterogenous effects of a massive content moderation intervention via Difference-in-Differences
- Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement Measurement
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- Exploring Content and Social Connections of Fake News with Explainable Text and Graph Learning
- Understanding the Interplay between LLMs' Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025