Generalizability vs. Robustness: Adversarial Examples for Medical Imaging
arXiv:1804.00504
Abstract
In this paper, for the first time, we propose an evaluation method for deep learning models that assesses the performance of a model not only in an unseen test scenario, but also in extreme cases of noise, outliers and ambiguous input data. To this end, we utilize adversarial examples, images that fool machine learning models, while looking imperceptibly different from original data, as a measure to evaluate the robustness of a variety of medical imaging models. Through extensive experiments on skin lesion classification and whole brain segmentation with state-of-the-art networks such as Inception and UNet, we show that models that achieve comparable performance regarding generalizability may have significant variations in their perception of the underlying data manifold, leading to an extensive performance gap in their robustness.
Under Review for MICCAI 2018
References in corpus (1)
Cited by in corpus (9)
- Adversarial Attack Vulnerability of Medical Image Analysis Systems: Unexplored Factors
- Adversarial Attacks Against Medical Deep Learning Systems
- Realistic Adversarial Data Augmentation for MR Image Segmentation
- Defending against adversarial attacks on medical imaging AI system, classification or detection?
- Adversarial Examples for Electrocardiograms
- No Surprises: Training Robust Lung Nodule Detection for Low-Dose CT Scans by Augmenting with Adversarial Attacks
- Enabling Data Diversity: Efficient Automatic Augmentation via Regularized Adversarial Training
- Adversarial Consistency and the Uniqueness of the Adversarial Bayes Classifier
- Tailored Bayes: a risk modelling framework under unequal misclassification costs