Multimodal Explainability via Latent Shift applied to COVID-19 stratification
arXiv:2212.14084 · doi:10.1016/j.patcog.2024.110825
Abstract
We are witnessing a widespread adoption of artificial intelligence in healthcare. However, most of the advancements in deep learning in this area consider only unimodal data, neglecting other modalities. Their multimodal interpretation necessary for supporting diagnosis, prognosis and treatment decisions. In this work we present a deep architecture, which jointly learns modality reconstructions and sample classifications using tabular and imaging data. The explanation of the decision taken is computed by applying a latent shift that, simulates a counterfactual prediction revealing the features of each modality that contribute the most to the decision and a quantitative score indicating the modality importance. We validate our approach in the context of COVID-19 pandemic using the AIforCOVID dataset, which contains multimodal data for the early identification of patients at risk of severe outcome. The results show that the proposed method provides meaningful explanations without degrading the classification performance.
References in corpus (3)
- A Deep Learning Approach for Virtual Contrast Enhancement in Contrast Enhanced Spectral Mammography
- A Deep Learning Approach for Overall Survival Prediction in Lung Cancer with Missing Values
- Multi-objective optimization determines when, which and how to fuse deep networks: an application to predict COVID-19 outcomes
Cited by in corpus (12)
- A Systematic Review of Intermediate Fusion in Multimodal Deep Learning for Biomedical Applications
- Class Balancing Diversity Multimodal Ensemble for Alzheimer's Disease Diagnosis and Early Detection
- MARIA: a Multimodal Transformer Model for Incomplete Healthcare Data
- A graph neural network-based model with Out-of-Distribution Robustness for enhancing Antiretroviral Therapy Outcome Prediction for HIV-1
- Multi-stage intermediate fusion for multimodal learning to classify non-small cell lung cancer subtypes from CT and PET
- Multi-Dataset Multi-Task Learning for COVID-19 Prognosis
- Whole-Body Image-to-Image Translation for a Virtual Scanner in a Healthcare Digital Twin
- Timing Is Everything: Finding the Optimal Fusion Points in Multimodal Medical Imaging
- Beyond a Single Mode: GAN Ensembles for Diverse Medical Data Generation
- Multimodal Doctor-in-the-Loop: A Clinically-Guided Explainable Framework for Predicting Pathological Response in Non-Small Cell Lung Cancer
- Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
- Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation