Multi-step retrieval and reasoning improves radiology question answering with large language models
arXiv:2508.00743 · doi:10.1038/s41746-025-02250-5
Abstract
Clinical decision-making in radiology increasingly benefits from artificial intelligence (AI), particularly through large language models (LLMs). However, traditional retrieval-augmented generation (RAG) systems for radiology question answering (QA) typically rely on single-step retrieval, limiting their ability to handle complex clinical reasoning tasks. Here we propose radiology Retrieval and Reasoning (RaR), a multi-step retrieval and reasoning framework designed to improve diagnostic accuracy, factual consistency, and clinical reliability of LLMs in radiology question answering. We evaluated 25 LLMs spanning diverse architectures, parameter scales (0.5B to >670B), and training paradigms (general-purpose, reasoning-optimized, clinically fine-tuned), using 104 expert-curated radiology questions from previously established RSNA-RadioQA and ExtendedQA datasets. To assess generalizability, we additionally tested on an unseen internal dataset of 65 real-world radiology board examination questions. RaR significantly improved mean diagnostic accuracy over zero-shot prompting and conventional online RAG. The greatest gains occurred in small-scale models, while very large models (>200B parameters) demonstrated minimal changes (<2% improvement). Additionally, RaR retrieval reduced hallucinations (mean 9.4%) and retrieved clinically relevant context in 46% of cases, substantially aiding factual grounding. Even clinically fine-tuned models showed gains from RaR (e.g., MedGemma-27B), indicating that retrieval remains beneficial despite embedded domain knowledge. These results highlight the potential of RaR to enhance factuality and diagnostic accuracy in radiology QA, warranting future studies to validate their clinical utility. All datasets, code, and the full RaR framework are publicly available to support open research and clinical translation.
Published in npj Digital Medicine
References in corpus (19)
- Survey of Hallucination in Natural Language Generation
- A Survey on Large Language Model based Autonomous Agents
- Scaling Laws for Neural Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V3 Technical Report
- Gemma: Open Models Based on Gemini Research and Technology
- Large Language Models Streamline Automated Machine Learning for Clinical Studies
- Qwen3 Technical Report
- Qwen Technical Report
- Gemma 3 Technical Report
- From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- MedGemma Technical Report
- RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
- Med42-v2: A Suite of Clinical LLMs
- MedVersa: A Generalist Foundation Model for Medical Image Interpretation
- Enhancing LLMs for Impression Generation in Radiology Reports through a Multi-Agent System
- CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering
- Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence
- Agent-Based Uncertainty Awareness Improves Automated Radiology Report Labeling with an Open-Source Large Language Model