Deep Weighted Averaging Classifiers
arXiv:1811.02579 · doi:10.1145/3287560.3287595
Abstract
Recent advances in deep learning have achieved impressive gains in classification accuracy on a variety of types of data, including images and text. Despite these gains, however, concerns have been raised about the calibration, robustness, and interpretability of these models. In this paper we propose a simple way to modify any conventional deep architecture to automatically provide more transparent explanations for classification decisions, as well as an intuitive notion of the credibility of each prediction. Specifically, we draw on ideas from nonparametric kernel regression, and propose to predict labels based on a weighted sum of training instances, where the weights are determined by distance in a learned instance-embedding space. Working within the framework of conformal methods, we propose a new measure of nonconformity suggested by our model, and experimentally validate the accompanying theoretical expectations, demonstrating improved transparency, controlled error rates, and robustness to out-of-domain data, without compromising on accuracy or calibration.
13 pages, 8 figures, 5 tables, added DOI and updated to meet ACM formatting requirements, In Proceedings of FAT* (2019)
References in corpus (19)
- Adam: A Method for Stochastic Optimization
- Explaining and Harnessing Adversarial Examples
- A Unified Approach to Interpreting Model Predictions
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Intriguing properties of neural networks
- Towards A Rigorous Science of Interpretable Machine Learning
- On Calibration of Modern Neural Networks
- A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
- A Sentimental Education: Sentiment Analysis Using Subjectivity Summarization Based on Minimum Cuts
- A tutorial on conformal prediction
- A Survey on Metric Learning for Feature Vectors and Structured Data
- Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning
- Grad-CAM: Why did you say that?
- Supersparse Linear Integer Models for Optimized Medical Scoring Systems
- To Trust Or Not To Trust A Classifier
- Interpreting Blackbox Models via Model Extraction
- The Bayesian Case Model: A Generative Approach for Case-Based Reasoning and Prototype Classification
- How do Humans Understand Explanations from Machine Learning Systems? An Evaluation of the Human-Interpretability of Explanation
- Conditional validity of inductive conformal predictors
Cited by in corpus (12)
- How Case Based Reasoning Explained Neural Networks: An XAI Survey of Post-Hoc Explanation-by-Example in ANN-CBR Twins
- Pretrained Transformers Improve Out-of-Distribution Robustness
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions
- Removing biased data to improve fairness and accuracy
- Wise-SrNet: A Novel Architecture for Enhancing Image Classification by Learning Spatial Resolution of Feature Maps
- Exemplar Auditing for Multi-Label Biomedical Text Classification
- Are Graph Neural Networks Miscalibrated?
- Deep Kernel Survival Analysis and Subject-Specific Survival Time Prediction Intervals
- Unsupervised Out-of-Domain Detection via Pre-trained Transformers
- SelfExplain: A Self-Explaining Architecture for Neural Text Classifiers
- Less is More: Rejecting Unreliable Reviews for Product Question Answering
- Influence Tuning: Demoting Spurious Correlations via Instance Attribution and Instance-Driven Updates