Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
arXiv:1903.12261
Abstract
In this paper we establish rigorous benchmarks for image classifier robustness. Our first benchmark, ImageNet-C, standardizes and expands the corruption robustness topic, while showing which classifiers are preferable in safety-critical applications. Then we propose a new dataset called ImageNet-P which enables researchers to benchmark a classifier's robustness to common perturbations. Unlike recent robustness research, this benchmark evaluates performance on common corruptions and perturbations not worst-case adversarial perturbations. We find that there are negligible changes in relative corruption robustness from AlexNet classifiers to ResNet classifiers. Afterward we discover ways to enhance corruption and perturbation robustness. We even find that a bypassed adversarial defense provides substantial common perturbation robustness. Together our benchmarks may aid future work toward networks that robustly generalize.
ICLR 2019 camera-ready; datasets available at https://github.com/hendrycks/robustness ; this article supersedes arXiv:1807.01697
Cited by in corpus (31)
- Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions
- RobustBench: a standardized adversarial robustness benchmark
- MEMO: Test Time Robustness via Adaptation and Augmentation
- Towards Evaluating the Robustness of Deep Diagnostic Models by Adversarial Attack
- Robustness Disparities in Commercial Face Detection
- Exposing Previously Undetectable Faults in Deep Neural Networks
- Blind Image Deblurring with Unknown Kernel Size and Substantial Noise
- Perfect density models cannot guarantee anomaly detection
- If a Human Can See It, So Should Your System: Reliability Requirements for Machine Vision Components
- How Good is Google Bard's Visual Understanding? An Empirical Study on Open Challenges
- On the effectiveness of adversarial training against common corruptions
- Infinitely Deep Bayesian Neural Networks with Stochastic Differential Equations
- CheXternal: Generalization of Deep Learning Models for Chest X-ray Interpretation to Photos of Chest X-rays and External Clinical Settings
- ACE: Adapting to Changing Environments for Semantic Segmentation
- Extending Environments To Measure Self-Reflection In Reinforcement Learning
- Deep Ensembles for Low-Data Transfer Learning
- Stationary Activations for Uncertainty Calibration in Deep Learning
- Denoised Internal Models: a Brain-Inspired Autoencoder against Adversarial Attacks
- Shifts: A Dataset of Real Distributional Shift Across Multiple Large-Scale Tasks
- CheXphotogenic: Generalization of Deep Learning Models for Chest X-ray Interpretation to Photos of Chest X-rays
- Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework
- A Self-Supervised Feature Map Augmentation (FMA) Loss and Combined Augmentations Finetuning to Efficiently Improve the Robustness of CNNs
- Compressive Visual Representations
- Uncover and Unlearn Nuisances: Agnostic Fully Test-Time Adaptation
- Minimal Sufficient Views: A DNN model making predictions with more evidence has higher accuracy
- Partial Wasserstein and Maximum Mean Discrepancy distances for bridging the gap between outlier detection and drift detection
- RBUE: A ReLU-Based Uncertainty Estimation Method of Deep Neural Networks
- Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
- Golden Grain: Building a Secure and Decentralized Model Marketplace for MLaaS
- Scene Uncertainty and the Wellington Posterior of Deterministic Image Classifiers
- LocalNorm: Robust Image Classification through Dynamically Regularized Normalization