On Calibration of Modern Neural Networks
arXiv:1706.04599
Abstract
Confidence calibration -- the problem of predicting probability estimates representative of the true correctness likelihood -- is important for classification models in many applications. We discover that modern neural networks, unlike those from a decade ago, are poorly calibrated. Through extensive experiments, we observe that depth, width, weight decay, and Batch Normalization are important factors influencing calibration. We evaluate the performance of various post-processing calibration methods on state-of-the-art architectures with image and document classification datasets. Our analysis and experiments not only offer insights into neural network learning, but also provide a simple and straightforward recipe for practical settings: on most datasets, temperature scaling -- a single-parameter variant of Platt Scaling -- is surprisingly effective at calibrating predictions.
ICML 2017
References in corpus (1)
Cited by in corpus (122)
- Physics-Constrained Deep Learning for High-dimensional Surrogate Modeling and Uncertainty Quantification without Labeled Data
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- SPINN: Synergistic Progressive Inference of Neural Networks over Device and Cloud
- Earthquake magnitude and location estimation from real time seismic waveforms with a transformer network
- Diminishing Uncertainty within the Training Pool: Active Learning for Medical Image Segmentation
- AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild
- Exploring the Limits of Out-of-Distribution Detection
- Adaptive Inference through Early-Exit Networks: Design, Challenges and Directions
- ExoMiner: A Highly Accurate and Explainable Deep Learning Classifier that Validates 301 New Exoplanets
- Human Activity Recognition on Microcontrollers with Quantized and Adaptive Deep Neural Networks
- Unsupervised Domain Adaptation of Black-Box Source Models
- Uncertainty Aware Training to Improve Deep Learning Model Calibration for Classification of Cardiac MR Images
- PAD: Towards Principled Adversarial Malware Detection Against Evasion Attacks
- Monitoring and explainability of models in production
- Trustable Co-label Learning from Multiple Noisy Annotators
- Adversarial Phenomenon in the Eyes of Bayesian Deep Learning
- Assaying Out-Of-Distribution Generalization in Transfer Learning
- Uncertainty-Aware Reliable Text Classification
- AI Research Considerations for Human Existential Safety (ARCHES)
- Toward Degree Bias in Embedding-Based Knowledge Graph Completion
- Role of Data Augmentation Strategies in Knowledge Distillation for Wearable Sensor Data
- Hybrid Discriminative-Generative Training via Contrastive Learning
- It's always personal: Using Early Exits for Efficient On-Device CNN Personalisation
- Wat zei je? Detecting Out-of-Distribution Translations with Variational Transformers
- Theoretical analysis and experimental validation of volume bias of soft Dice optimized segmentation maps in the context of inherent uncertainty
- Cal-SFDA: Source-Free Domain-adaptive Semantic Segmentation with Differentiable Expected Calibration Error
- Set-valued classification -- overview via a unified framework
- Early Rumor Detection Using Neural Hawkes Process with a New Benchmark Dataset
- An Empirical Analysis of the Impact of Data Augmentation on Knowledge Distillation
- Learning Uncertainty For Safety-Oriented Semantic Segmentation In Autonomous Driving
- Diverse Ensembles Improve Calibration
- Opportunistic Learning: Budgeted Cost-Sensitive Learning from Data Streams
- Acoustic scene classification using teacher-student learning with soft-labels
- Detecting, Localising and Classifying Polyps from Colonoscopy Videos using Deep Learning
- Are all outliers alike? On Understanding the Diversity of Outliers for Detecting OODs
- ReMix: Calibrated Resampling for Class Imbalance in Deep learning
- Should Ensemble Members Be Calibrated?
- Active Speaker Detection as a Multi-Objective Optimization with Uncertainty-based Multimodal Fusion
- Early-stopped neural networks are consistent
- Assurance Monitoring of Learning Enabled Cyber-Physical Systems Using Inductive Conformal Prediction based on Distance Learning
- Bayesian Graph Neural Networks for Molecular Property Prediction
- Unsupervised Domain Adaptation in the Absence of Source Data
- Interpretable Uncertainty Quantification in AI for HEP
- DeepLens: Interactive Out-of-distribution Data Detection in NLP Models
- Leveraging Angular Distributions for Improved Knowledge Distillation
- Change-point detection in anomalous-diffusion trajectories utilising machine-learning-based uncertainty estimates
- Detecting OODs as datapoints with High Uncertainty
- Controlling Risk of Web Question Answering
- Human Evaluation of Spoken vs. Visual Explanations for Open-Domain QA
- On the Role of Dataset Quality and Heterogeneity in Model Confidence
- Semi-Supervised Learning with Normalizing Flows
- On the use of uncertainty in classifying Aedes Albopictus mosquitoes
- Probing the Purview of Neural Networks via Gradient Analysis
- Deep Active Learning for Efficient Training of a LiDAR 3D Object Detector
- Closer Look at the Uncertainty Estimation in Semantic Segmentation under Distributional Shift
- The Battleship Approach to the Low Resource Entity Matching Problem
- Interpreting Neural Networks Using Flip Points
- Improving Robustness to Model Inversion Attacks via Mutual Information Regularization
- Consistent Sparse Deep Learning: Theory and Computation
- PAC Confidence Predictions for Deep Neural Network Classifiers
- Ranking over Regression for Bayesian Optimization and Molecule Selection
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- A Comparative Study of Calibration Methods for Imbalanced Class Incremental Learning
- Robust Reading Comprehension with Linguistic Constraints via Posterior Regularization
- Beyond Point Estimate: Inferring Ensemble Prediction Variation from Neuron Activation Strength in Recommender Systems
- Spatially Varying Label Smoothing: Capturing Uncertainty from Expert Annotations
- A Targeted Universal Attack on Graph Convolutional Network
- Unsupervised Temperature Scaling: An Unsupervised Post-Processing Calibration Method of Deep Networks
- Robust quantum dots charge autotuning using neural network uncertainty
- Intelligence plays dice: Stochasticity is essential for machine learning
- Progressively Complementary Network for Fisheye Image Rectification Using Appearance Flow
- Energy-Based Models with Applications to Speech and Language Processing
- Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowledge Distillation
- Improving Classifier Confidence using Lossy Label-Invariant Transformations
- Exploiting Large-scale Teacher-Student Training for On-device Acoustic Models
- Learning Domain Adaptation with Model Calibration for Surgical Report Generation in Robotic Surgery
- On Modelling Label Uncertainty in Deep Neural Networks: Automatic Estimation of Intra-observer Variability in 2D Echocardiography Quality Assessment
- On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective
- Deep Transfer Learning with Ridge Regression
- Ask-n-Learn: Active Learning via Reliable Gradient Representations for Image Classification
- Calibration and Uncertainty for multiRater Volume Assessment in multiorgan Segmentation (CURVAS) challenge results
- The Misclassification of Autistic Writing as AI-Generated
- Improving Uncertainty-Error Correspondence in Deep Bayesian Medical Image Segmentation
- On the Calibration and Uncertainty of Neural Learning to Rank Models
- Towards User Guided Actionable Recourse
- Estimating Predictive Uncertainty Under Program Data Distribution Shift
- A witness function based construction of discriminative models using Hermite polynomials
- Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection In Task Oriented Dialog
- Learning ULMFiT and Self-Distillation with Calibration for Medical Dialogue System
- Image-Based Jet Analysis
- RBUE: A ReLU-Based Uncertainty Estimation Method of Deep Neural Networks
- Diversifying Dialog Generation via Adaptive Label Smoothing
- Power Pooling Operators and Confidence Learning for Semi-Supervised Sound Event Detection
- Trinary Tools for Continuously Valued Binary Classifiers
- Doppler Spectrum Classification with CNNs via Heatmap Location Encoding and a Multi-head Output Layer
- Uncertainty Propagation in Node Classification
- Online Black-Box Confidence Estimation of Deep Neural Networks
- Performance Measurement for Deep Bayesian Neural Network
- MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
- Deep inference of simulated strong lenses in ground-based surveys
- Machine-learned trends in mirror configurations in the Large Plasma Device
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift
- An analysis of gamma-ray data collected at traffic intersections in Northern Virginia
- Know Where To Drop Your Weights: Towards Faster Uncertainty Estimation
- Long Short-Term Sample Distillation
- Self Training with Ensemble of Teacher Models
- Right Decisions from Wrong Predictions: A Mechanism Design Alternative to Individual Calibration
- The Robust Semantic Segmentation UNCV2023 Challenge Results
- Multi-Sample Online Learning for Spiking Neural Networks based on Generalized Expectation Maximization
- On Class Imbalance and Background Filtering in Visual Relationship Detection
- Natural Attribute-based Shift Detection
- Teaching Uncertainty Quantification in Machine Learning through Use Cases
- Confidence Preservation Property in Knowledge Distillation Abstractions
- MACEst: The reliable and trustworthy Model Agnostic Confidence Estimator
- Assessing The Importance Of Colours For CNNs In Object Recognition
- Bayesian deep learning of affordances from RGB images
- Generalization by Recognizing Confusion
- Understanding Classifier Mistakes with Generative Models
- MIMIR: Deep Regression for Automated Analysis of UK Biobank Body MRI
- Pathologies in priors and inference for Bayesian transformers
- UNGOML: Automated Classification of unsafe Usages in Go
- ALT-MAS: A Data-Efficient Framework for Active Testing of Machine Learning Algorithms