Multi-objective optimization determines when, which and how to fuse deep networks: an application to predict COVID-19 outcomes
arXiv:2204.03772 · doi:10.1016/j.compbiomed.2023.106625
Abstract
The COVID-19 pandemic has caused millions of cases and deaths and the AI-related scientific community, after being involved with detecting COVID-19 signs in medical images, has been now directing the efforts towards the development of methods that can predict the progression of the disease. This task is multimodal by its very nature and, recently, baseline results achieved on the publicly available AIforCOVID dataset have shown that chest X-ray scans and clinical information are useful to identify patients at risk of severe outcomes. While deep learning has shown superior performance in several medical fields, in most of the cases it considers unimodal data only. In this respect, when, which and how to fuse the different modalities is an open challenge in multimodal deep learning. To cope with these three questions here we present a novel approach optimizing the setup of a multimodal end-to-end model. It exploits Pareto multi-objective optimization working with a performance metric and the diversity score of multiple candidate unimodal neural networks to be fused. We test our method on the AIforCOVID dataset, attaining state-of-the-art results, not only outperforming the baseline performance but also being robust to external validation. Moreover, exploiting XAI algorithms we figure out a hierarchy among the modalities and we extract the features' intra-modality importance, enriching the trust on the predictions made by the model.
References in corpus (3)
Cited by in corpus (12)
- A Systematic Review of Intermediate Fusion in Multimodal Deep Learning for Biomedical Applications
- Multimodal Explainability via Latent Shift applied to COVID-19 stratification
- Class Balancing Diversity Multimodal Ensemble for Alzheimer's Disease Diagnosis and Early Detection
- MARIA: a Multimodal Transformer Model for Incomplete Healthcare Data
- A graph neural network-based model with Out-of-Distribution Robustness for enhancing Antiretroviral Therapy Outcome Prediction for HIV-1
- Multi-Dataset Multi-Task Learning for COVID-19 Prognosis
- Whole-Body Image-to-Image Translation for a Virtual Scanner in a Healthcare Digital Twin
- Timing Is Everything: Finding the Optimal Fusion Points in Multimodal Medical Imaging
- Beyond a Single Mode: GAN Ensembles for Diverse Medical Data Generation
- Multimodal Doctor-in-the-Loop: A Clinically-Guided Explainable Framework for Predicting Pathological Response in Non-Small Cell Lung Cancer
- Evaluating Vision Language Model Adaptations for Radiology Report Generation in Low-Resource Languages
- Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation