Survey of Hallucination in Natural Language Generation
arXiv:2202.03629 · doi:10.1145/3571730
Abstract
Natural Language Generation (NLG) has improved exponentially in recent years thanks to the development of sequence-to-sequence deep learning technologies such as Transformer-based language models. This advancement has led to more fluent and coherent NLG, leading to improved development in downstream tasks such as abstractive summarization, dialogue generation and data-to-text generation. However, it is also apparent that deep learning based generation is prone to hallucinate unintended text, which degrades the system performance and fails to meet user expectations in many real-world scenarios. To address this issue, many studies have been presented in measuring and mitigating hallucinated texts, but these have never been reviewed in a comprehensive manner before. In this survey, we thus provide a broad overview of the research progress and challenges in the hallucination problem in NLG. The survey is organized into two parts: (1) a general overview of metrics, mitigation methods, and future directions; (2) an overview of task-specific research progress on hallucinations in the following downstream tasks, namely abstractive summarization, dialogue generation, generative question answering, data-to-text generation, machine translation, and visual-language generation; and (3) hallucinations in large language models (LLMs). This survey serves to facilitate collaborative efforts among researchers in tackling the challenge of hallucinated texts in NLG.
References in corpus (15)
- Flamingo: a Visual Language Model for Few-Shot Learning
- Data Augmentation Approaches in Natural Language Processing: A Survey
- Survey on reinforcement learning for language processing
- The Factual Inconsistency Problem in Abstractive Text Summarization: A Survey
- Factuality Enhanced Language Models for Open-Ended Text Generation
- Open-Domain Conversational Agents: Current Progress, Open Problems, and Future Directions
- CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning
- Revisiting Challenges in Data-to-Text Generation with Fact Grounding
- Global-to-local Memory Pointer Networks for Task-Oriented Dialogue
- WebGPT: Browser-assisted question-answering with human feedback
- Deduplicating Training Data Makes Language Models Better
- SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization
- Rome was built in 1776: A Case Study on Factual Correctness in Knowledge-Grounded Response Generation
- Hallucination of speech recognition errors with sequence to sequence learning
- DialFact: A Benchmark for Fact-Checking in Dialogue
Cited by in corpus (188)
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- A Survey on Large Language Model based Autonomous Agents
- Generative AI
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- Recommender Systems in the Era of Large Language Models (LLMs)
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- The Robots are Here: Navigating the Generative AI Revolution in Computing Education
- Unleashing the potential of prompt engineering for large language models
- Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
- The Metacognitive Demands and Opportunities of Generative AI
- GenAI Against Humanity: Nefarious Applications of Generative Artificial Intelligence and Large Language Models
- Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
- GPT Models in Construction Industry: Opportunities, Limitations, and a Use Case Validation
- AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
- How Generative AI models such as ChatGPT can be (Mis)Used in SPC Practice, Education, and Research? An Exploratory Study
- Transformers in Healthcare: A Survey
- GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information
- Factuality Challenges in the Era of Large Language Models
- Generative AI in the Construction Industry: Opportunities & Challenges
- Understanding Large-Language Model (LLM)-powered Human-Robot Interaction
- Delving into LLM-assisted writing in biomedical publications through excess vocabulary
- "I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust
- Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education
- Harms from Increasingly Agentic Algorithmic Systems
- MindfulDiary: Harnessing Large Language Model to Support Psychiatric Patients' Journaling
- How understanding large language models can inform the use of ChatGPT in physics education
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- Mathemyths: Leveraging Large Language Models to Teach Mathematical Language through Child-AI Co-Creative Storytelling
- Tool Learning with Large Language Models: A Survey
- General Purpose Artificial Intelligence Systems (GPAIS): Properties, Definition, Taxonomy, Societal Implications and Responsible Governance
- Deception Abilities Emerged in Large Language Models
- From COBIT to ISO 42001: Evaluating Cybersecurity Frameworks for Opportunities, Risks, and Regulatory Compliance in Commercializing Large Language Models
- Careless Whisper: Speech-to-Text Hallucination Harms
- Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
- Black-Box Access is Insufficient for Rigorous AI Audits
- Abstractive Text Summarization: State of the Art, Challenges, and Improvements
- HILL: A Hallucination Identifier for Large Language Models
- Retrieving Supporting Evidence for Generative Question Answering
- A Practical Survey on Zero-shot Prompt Design for In-context Learning
- Automated Educational Question Generation at Different Bloom's Skill Levels using Large Language Models: Strategies and Evaluation
- Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
- The illusion of artificial inclusion
- Knowledge graph enhanced retrieval-augmented generation for failure mode and effects analysis
- The Great AI Witch Hunt: Reviewers Perception and (Mis)Conception of Generative AI in Research Writing
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
- Understanding Users' Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level
- When Geoscience Meets Generative AI and Large Language Models: Foundations, Trends, and Future Challenges
- ChatGPT in Veterinary Medicine: A Practical Guidance of Generative Artificial Intelligence in Clinics, Education, and Research
- ABScribe: Rapid Exploration & Organization of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Language Models
- Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
- Large Language Model for Table Processing: A Survey
- Application of NotebookLM, a Large Language Model with Retrieval-Augmented Generation, for Lung Cancer Staging
- Evaluating Generative Ad Hoc Information Retrieval
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
- MindShift: Leveraging Large Language Models for Mental-States-Based Problematic Smartphone Use Intervention
- Relatedly: Scaffolding Literature Reviews with Existing Related Work Sections
- KernelGPT: Enhanced Kernel Fuzzing via Large Language Models
- Can GPT-3.5 Generate and Code Discharge Summaries?
- Cross-Data Knowledge Graph Construction for LLM-enabled Educational Question-Answering System: A Case Study at HCMUT
- One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
- QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
- Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
- Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
- Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization
- Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data
- "It's like a rubber duck that talks back": Understanding Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study
- RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
- Deception and Manipulation in Generative AI
- DOPRA: Decoding Over-accumulation Penalization and Re-allocation in Specific Weighting Layer
- Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing Scenarios
- Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
- A Reliable Knowledge Processing Framework for Combustion Science using Foundation Models
- From human-centered to social-centered artificial intelligence: Assessing ChatGPT's impact through disruptive events
- Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
- Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models
- Collage is the New Writing: Exploring the Fragmentation of Text and User Interfaces in AI Tools
- Specification Overfitting in Artificial Intelligence
- Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
- Investigating Interaction Modes and User Agency in Human-LLM Collaboration for Domain-Specific Data Analysis
- Explainability for Transparent Conversational Information-Seeking
- VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
- Computational Argumentation-based Chatbots: a Survey
- On Early Detection of Hallucinations in Factual Question Answering
- Bridging Generations using AI-Supported Co-Creative Activities
- Enhancing Tourism Recommender Systems for Sustainable City Trips Using Retrieval-Augmented Generation
- Examining risks of racial biases in NLP tools for child protective services
- ShennongAlpha: an AI-driven sharing and collaboration platform for intelligent curation, acquisition, and translation of natural medicinal material knowledge
- Multi-step retrieval and reasoning improves radiology question answering with large language models
- Unlocking the Potential of Past Research: Using Generative AI to Reconstruct Healthcare Simulation Models
- Narrative Player: Reviving Data Narratives with Visuals
- Learned, uncertainty-driven adaptive acquisition for photon-efficient scanning microscopy
- A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
- Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
- Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code Completions
- The Responsible Development of Automated Student Feedback with Generative AI
- Prompt Injection Attacks in Defended Systems
- Complex System Diagnostics Using a Knowledge Graph-Informed and Large Language Model-Enhanced Framework
- Walert: Putting Conversational Search Knowledge into Action by Building and Evaluating a Large Language Model-Powered Chatbot
- Generative artificial intelligence in dentistry: Current approaches and future challenges
- Towards Filling the Gap in Conversational Search: From Passage Retrieval to Conversational Response Generation
- An empathic GPT-based chatbot to talk about mental disorders with Spanish teenagers
- A Knowledge-Informed Deep Learning Paradigm for Generalizable and Stability-Optimized Car-Following Models
- Symbolic Equation Solving via Reinforcement Learning
- SPROUT: an Interactive Authoring Tool for Generating Programming Tutorials with the Visualization of Large Language Models
- Consolidating Trees of Robotic Plans Generated Using Large Language Models to Improve Reliability
- Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
- Normative Conflicts and Shallow AI Alignment
- Leveraging Large Language Models through Natural Language Processing to provide interpretable Machine Learning predictions of mental deterioration in real time
- A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language Models
- A Research Roadmap for Augmenting Software Engineering Processes and Software Products with Generative AI
- Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
- LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
- GPT Assisted Annotation of Rhetorical and Linguistic Features for Interpretable Propaganda Technique Detection in News Text
- A Hybrid Intelligence Method for Argument Mining
- A Survey on Self-Supervised Graph Foundation Models: Knowledge-Based Perspective
- Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models
- Can GPT-4o Evaluate Usability Like Human Experts? A Comparative Study on Issue Identification in Heuristic Evaluation
- Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
- Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
- Evaluating LLMs for Visualization Generation and Understanding
- SciDaSynth: Interactive Structured Data Extraction from Scientific Literature with Large Language Model
- Complex QA and language models hybrid architectures, Survey
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- ExClaim: Explainable Neural Claim Verification Using Rationalization
- TrICy: Trigger-guided Data-to-text Generation with Intent aware Attention-Copy
- Vision-Based Hand Gesture Customization from a Single Demonstration
- Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking
- Knowledge Graphs for Enhancing Large Language Models in Entity Disambiguation
- IOAgent: Democratizing Trustworthy HPC I/O Performance Diagnosis Capability via LLMs
- Unlocking Electronic Health Records: A Hybrid Graph RAG Approach to Safe Clinical AI for Patient QA
- Behind the Counter: Exploring the Motivations and Barriers of Online Counterspeech Writing
- generAItor: Tree-in-the-Loop Text Generation for Language Model Explainability and Adaptation
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- ProSLM : A Prolog Synergized Language Model for explainable Domain Specific Knowledge Based Question Answering
- Improving Public Service Chatbot Design and Civic Impact: Investigation of Citizens' Perceptions of a Metro City 311 Chatbot
- Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
- GRILLBot In Practice: Lessons and Tradeoffs Deploying Large Language Models for Adaptable Conversational Task Assistants
- Generative AI for Research Data Processing: Lessons Learnt From Three Use Cases
- The Importance of Distrust in AI
- Can Users Detect Biases or Factual Errors in Generated Responses in Conversational Information-Seeking?
- CACER: Clinical Concept Annotations for Cancer Events and Relations
- Information Retrieval in the Age of Generative AI: The RGB Model
- Highlighting Case Studies in LLM Literature Review of Interdisciplinary System Science
- Building the ethical AI framework of the future: from philosophy to practice
- PLAID: Supporting Computing Instructors to Identify Domain-Specific Programming Plans at Scale
- Adversarial Text Rewriting for Text-aware Recommender Systems
- Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
- RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
- Visual hallucination detection in large vision-language models via evidential conflict
- Towards Self-Contained Answers: Entity-Based Answer Rewriting in Conversational Search
- AI and Agile Software Development: From Frustration to Success -- XP2025 Workshop Summary
- "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking Partner
- LLM-Confidence Reranker: A Training-Free Approach for Enhancing Retrieval-Augmented Generation Systems
- Efficient Serving of LLM Applications with Probabilistic Demand Modeling
- AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
- AgriIR: A Scalable Framework for Domain-Specific Knowledge Retrieval
- Align Generative Artificial Intelligence with Human Preferences: A Novel Large Language Model Fine-Tuning Method for Online Review Management
- When AI reviews science: Can we trust the referee?
- Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
- Medication counseling with large language models: balancing flexibility and rigidity
- iJTyper: An Iterative Type Inference Framework for Java by Integrating Constraint- and Statistically-based Methods
- ACTS: A multi-tier benchmark evaluating LLM cipher identification under controlled blind conditions
- NoTeeline: Supporting Real-Time, Personalized Notetaking with LLM-Enhanced Micronotes
- LP-LM: No Hallucinations in Question Answering with Logic Programming
- LLM-PQA: LLM-enhanced Prediction Query Answering
- Clean Up the Mess: Addressing Data Pollution in Cryptocurrency Abuse Reporting Services
- A Roadmap for Tamed Interactions with Large Language Models
- Online Domain-aware LLM Decoding for Continual Domain Evolution
- A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
- When AI Evaluates Its Own Work: Validating Learner-Initiated, AI-Generated Physics Practice Problems
- Trust-Aware Routing for Distributed Generative AI Inference at the Edge
- Spatial Priming Outperforms Semantic Prompting: A Grid-Based Approach to Improving LLM Accuracy on Chart Data Extraction
- Derivation Prompting: A Logic-Based Method for Improving Retrieval-Augmented Generation
- Short paper: Models in the dark -- Rectification and erasure under GDPR in ML supply chains
- Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation
- When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models
- CT-SAFR: Safe and Interpretable Chain-of-Thought Reasoning for Autonomous Robots: A Multi-Layered Verification Framework for Trustworthy AI-Driven Robotic Decision Making
- Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
- Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers?
- Hallucination Detection in Large Language Models Using Diversion Decoding
- Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning
- From Trait to Behavior: A Cognitive-Affective Personality System (CAPS) Perspective on Multi-Homing Intention in AIGC Platforms
- RoboCritics: Enabling Reliable End-to-End LLM Robot Programming through Expert-Informed Critics
- Towards Robust Retrieval-Augmented Generation Based on Knowledge Graph: A Comparative Analysis
- BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
- Generating Faithful Text From a Knowledge Graph with Noisy Reference Text