Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
arXiv:2608.02238 · doi:10.1088/2516-1091/ae4e74
Abstract
Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital health. Key dimensions, including robustness, explainability, fairness, accountability, and privacy, need to be addressed throughout the AI lifecycle, from problem formulation and data collection to model deployment and human interaction. While various contributions address different aspects of trustworthy AI, a focused synthesis on robustness and explainability, especially tailored to the healthcare context, remains limited. This review addresses that need by organizing recent advancements into an accessible framework, highlighting both technical and practical considerations. We present a structured overview of methods, challenges, and solutions, aiming to support researchers and practitioners in developing reliable and explainable AI solutions for digital health. This review article is organized into three main parts. First, we introduce the pillars of trustworthy AI and discuss the technical and ethical challenges, particularly in the context of digital health. Second, we explore application-specific trust considerations across domains such as intensive care, neonatal health, and metabolic health, highlighting how robustness and explainability support trust. Lastly, we present recent advancements in techniques aimed at improving robustness under data scarcity and distributional shifts, as well as explainable AI methods ranging from feature attribution to gradient-based interpretations and counterfactual explanations. This paper is further enriched with detailed discussions of the contributions toward robustness and explainability in digital health, the development of trustworthy AI systems in the era of LLMs, and various evaluation metrics for measuring trust and related parameters such as validity, fidelity, and diversity.
Preprint of the paper published in Progress in Biomedical Engineering. 26 pages, 5 figures
References in corpus (11)
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Domain Adaptation for Medical Image Analysis: A Survey
- Fairness And Bias in Artificial Intelligence: A Brief Survey of Sources, Impacts, And Mitigation Strategies
- From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
- Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation
- The Evolution of Distributed Systems for Graph Neural Networks and their Origin in Graph Processing and Deep Learning: A Survey
- Sample Selection Bias in Machine Learning for Healthcare
- LLM-Powered Prediction of Hyperglycemia and Discovery of Behavioral Treatment Pathways from Wearables and Diet
- RNAS-CL: Robust Neural Architecture Search by Cross-Layer Knowledge Distillation
- Inter-Beat Interval Estimation with Tiramisu Model: A Novel Approach with Reduced Error
- Enhancing Metabolic Syndrome Prediction with Hybrid Data Balancing and Counterfactuals