Prototype selection for interpretable classification
arXiv:1202.5933 · doi:10.1214/11-AOAS495
Abstract
Prototype methods seek a minimal subset of samples that can serve as a distillation or condensed view of a data set. As the size of modern data sets grows, being able to present a domain specialist with a short list of "representative" samples chosen from the data set is of increasing interpretative value. While much recent statistical research has been focused on producing sparse-in-the-variables methods, this paper aims at achieving sparsity in the samples. We discuss a method for selecting prototypes in the classification setting (in which the samples fall into known discrete categories). Our method of focus is derived from three basic properties that we believe a good prototype set should satisfy. This intuition is translated into a set cover optimization problem, which we solve approximately using standard approaches. While prototype selection is usually viewed as purely a means toward building an efficient classifier, in this paper we emphasize the inherent value of having a set of prototypical elements. That said, by using the nearest-neighbor rule on the set of prototypes, we can of course discuss our method as a classifier as well.
Published in at http://dx.doi.org/10.1214/11-AOAS495 the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org). arXiv admin note: text overlap with arXiv:0908.2284
Cited by in corpus (10)
- Methods and Models for Interpretable Linear Classification
- 10,000+ Times Accelerated Robust Subset Selection (ARSS)
- Distribution Density, Tails, and Outliers in Machine Learning: Metrics and Applications
- Melody: Generating and Visualizing Machine Learning Model Summary to Understand Data and Classifiers Together
- SLEEPER: interpretable Sleep staging via Prototypes from Expert Rules
- Learning a sparse database for patch-based medical image segmentation
- Safety design concepts for statistical machine learning components toward accordance with functional safety standards
- Unifying Model Explainability and Robustness via Machine-Checkable Concepts
- Enumeration of Distinct Support Vectors for Interactive Decision Making
- Interpretable Multiple-Kernel Prototype Learning for Discriminative Representation and Feature Selection