Utilizing Network Properties to Detect Erroneous Inputs
arXiv:2002.12520
Abstract
Neural networks are vulnerable to a wide range of erroneous inputs such as adversarial, corrupted, out-of-distribution, and misclassified examples. In this work, we train a linear SVM classifier to detect these four types of erroneous data using hidden and softmax feature vectors of pre-trained neural networks. Our results indicate that these faulty data types generally exhibit linearly separable activation properties from correct examples, giving us the ability to reject bad inputs with no extra training or overhead. We experimentally validate our findings across a diverse range of datasets, domains, pre-trained models, and adversarial attacks.
References in corpus (12)
- Generalisation in humans and deep neural networks
- Learning Confidence for Out-of-Distribution Detection in Neural Networks
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
- Natural Adversarial Examples
- Selective Classification for Deep Neural Networks
- Adversarial vulnerability for any classifier
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
- Training Confidence-calibrated Classifiers for Detecting Out-of-Distribution Samples
- Open Category Detection with PAC Guarantees
- Discriminative out-of-distribution detection for semantic segmentation
- Improved Adversarial Robustness by Reducing Open Space Risk via Tent Activations