A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination
arXiv:2505.00008 · doi:10.1016/j.jbi.2025.104866
Abstract
Objective: This review aims to explore the potential and challenges of using Natural Language Processing (NLP) to detect, correct, and mitigate medically inaccurate information, including errors, misinformation, and hallucination. By unifying these concepts, the review emphasizes their shared methodological foundations and their distinct implications for healthcare. Our goal is to advance patient safety, improve public health communication, and support the development of more reliable and transparent NLP applications in healthcare. Methods: A scoping review was conducted following PRISMA guidelines, analyzing studies from 2020 to 2024 across five databases. Studies were selected based on their use of NLP to address medically inaccurate information and were categorized by topic, tasks, document types, datasets, models, and evaluation metrics. Results: NLP has shown potential in addressing medically inaccurate information on the following tasks: (1) error detection (2) error correction (3) misinformation detection (4) misinformation correction (5) hallucination detection (6) hallucination mitigation. However, challenges remain with data privacy, context dependency, and evaluation standards. Conclusion: This review highlights the advancements in applying NLP to tackle medically inaccurate information while underscoring the need to address persistent challenges. Future efforts should focus on developing real-world datasets, refining contextual methods, and improving hallucination management to ensure reliable and transparent healthcare applications.
This paper has been accepted by the Journal of Biomedical Informatics
References in corpus (18)
- Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- Artificial Intelligence for Health Message Generation: Theory, Method, and an Empirical Study Using Prompt Engineering
- Can large language models reason about medical questions?
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- DR.BENCH: Diagnostic Reasoning Benchmark for Clinical Natural Language Processing
- Claim Detection for Automated Fact-checking: A Survey on Monolingual, Multilingual and Cross-Lingual Research
- Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
- Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations
- Med-MMHL: A Multi-Modal Dataset for Detecting Human- and LLM-Generated Misinformation in the Medical Domain
- MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
- Entity-based Claim Representation Improves Fact-Checking of Medical Content in Tweets
- Large Language Models in the Clinic: A Comprehensive Benchmark
- Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
- MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context
- Minimizing Factual Inconsistency and Hallucination in Large Language Models
- Listening to Patients: A Framework of Detecting and Mitigating Patient Misreport for Medical Dialogue Generation