artificial intelligence in medicine

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

arXiv:2607.25489

summary

The paper reviews how large language and multimodal models are being used as autonomous agents in medical tasks, covering their architectures, applications, evaluation methods, and the challenges for moving these systems into real clinical practice.

Abstract

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

Review article, 6 figures, 2 tables. 42 pages

Topics & keywords

#agentic ai#clinical decision support#multimodal models#evaluation#medical imaging#workflow integrationlarge language modelsmultimodal foundation modelstool usemulti‑agent systemsclinical translationprospective validationauditability