On the Robustness of Pretraining and Self-Supervision for a Deep Learning-based Analysis of Diabetic Retinopathy
arXiv:2106.13497
Abstract
There is an increasing number of medical use-cases where classification algorithms based on deep neural networks reach performance levels that are competitive with human medical experts. To alleviate the challenges of small dataset sizes, these systems often rely on pretraining. In this work, we aim to assess the broader implications of these approaches. For diabetic retinopathy grading as exemplary use case, we compare the impact of different training procedures including recently established self-supervised pretraining methods based on contrastive learning. To this end, we investigate different aspects such as quantitative performance, statistics of the learned feature representations, interpretability and robustness to image distortions. Our results indicate that models initialized from ImageNet pretraining report a significant increase in performance, generalization and robustness to image distortions. In particular, self-supervised models show further benefits to supervised models. Self-supervised models with initialization from ImageNet pretraining not only report higher performance, they also reduce overfitting to large lesions along with improvements in taking into account minute lesions indicative of the progression of the disease. Understanding the effects of pretraining in a broader sense that goes beyond simple performance comparisons is of crucial importance for the broader medical imaging community beyond the use-case considered in this work.
References in corpus (6)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Methods for Interpreting and Understanding Deep Neural Networks
- Med3D: Transfer Learning for 3D Medical Image Analysis
- A Systematic Comparison of Bayesian Deep Learning Robustness in Diabetic Retinopathy Tasks
- COVID-19 Prognosis via Self-Supervised Representation Learning and Multi-Image Prediction
- The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models