Supersparse Linear Integer Models for Optimized Medical Scoring Systems
arXiv:1502.04269 · doi:10.1007/s10994-015-5528-6
Abstract
Scoring systems are linear classification models that only require users to add, subtract and multiply a few small numbers in order to make a prediction. These models are in widespread use by the medical community, but are difficult to learn from data because they need to be accurate and sparse, have coprime integer coefficients, and satisfy multiple operational constraints. We present a new method for creating data-driven scoring systems called a Supersparse Linear Integer Model (SLIM). SLIM scoring systems are built by solving an integer program that directly encodes measures of accuracy (the 0-1 loss) and sparsity (the -seminorm) while restricting coefficients to coprime integers. SLIM can seamlessly incorporate a wide range of operational constraints related to accuracy and sparsity, and can produce highly tailored models without parameter tuning. We provide bounds on the testing and training accuracy of SLIM scoring systems, and present a new data reduction technique that can improve scalability by eliminating a portion of the training data beforehand. Our paper includes results from a collaboration with the Massachusetts General Hospital Sleep Laboratory, where SLIM was used to create a highly tailored scoring system for sleep apnea screening
This version reflects our findings on SLIM as of January 2016 (arXiv:1306.5860 and arXiv:1405.4047 are out-of-date). The final published version of this articled is available at http://www.springerlink.com
Cited by in corpus (71)
- A Survey on the Explainability of Supervised Machine Learning
- Model-Agnostic Interpretability of Machine Learning
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Actionable Recourse in Linear Classification
- Fairness and Accountability Design Needs for Algorithmic Support in High-Stakes Public Sector Decision-Making
- "Why Should I Trust You?": Explaining the Predictions of Any Classifier
- What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use
- Deep learning for temporal data representation in electronic health records: A systematic review of challenges and methodologies
- Interpretable Classification Models for Recidivism Prediction
- Interpreting Blackbox Models via Model Extraction
- An Evaluation of the Human-Interpretability of Explanation
- Learning Certifiably Optimal Rule Lists for Categorical Data
- How do Humans Understand Explanations from Machine Learning Systems? An Evaluation of the Human-Interpretability of Explanation
- Interpretability via Model Extraction
- On the Existence of Simpler Machine Learning Models
- The Road to Explainability is Paved with Bias: Measuring the Fairness of Explanations
- Quantifying Model Complexity via Functional Decomposition for Better Post-Hoc Interpretability
- An Interpretable Model with Globally Consistent Explanations for Credit Risk
- A study on the Interpretability of Neural Retrieval Models using DeepSHAP
- Sparsity in Optimal Randomized Classification Trees
- Assessing the Local Interpretability of Machine Learning Models
- Optimal randomized classification trees
- Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead
- Deep Weighted Averaging Classifiers
- Programs as Black-Box Explanations
- Predictive Multiplicity in Classification
- Confounding-Robust Policy Improvement
- Explanations of Black-Box Model Predictions by Contextual Importance and Utility
- AutoScore-Survival: Developing interpretable machine learning-based time-to-event scores with right-censored survival data
- Model Distillation for Revenue Optimization: Interpretable Personalized Pricing
- The Price of Interpretability
- On the Art and Science of Machine Learning Explanations
- Knowledge-based Transfer Learning Explanation
- A Semidefinite Programming Method for Integer Convex Quadratic Minimization
- Or's of And's for Interpretable Classification, with Application to Context-Aware Recommender Systems
- Quadratic Convergence of Smoothing Newton's Method for 0/1 Loss Optimization
- Bias in Data-driven AI Systems -- An Introductory Survey
- Embedding Deep Networks into Visual Explanations
- Model extraction from counterfactual explanations
- The Authority of "Fair" in Machine Learning
- Explainable AI using expressive Boolean formulas
- A Decision-Theoretic Approach for Model Interpretability in Bayesian Framework
- Proposed Guidelines for the Responsible Use of Explainable Machine Learning
- Learning Optimized Or's of And's
- FedScore: A privacy-preserving framework for federated scoring system development
- Learning Interpretable Models with Causal Guarantees
- VINE: Visualizing Statistical Interactions in Black Box Models
- Hybrid Predictive Model: When an Interpretable Model Collaborates with a Black-box Model
- Improving the Interpretability of Deep Neural Networks with Knowledge Distillation
- A Holistic Approach to Interpretability in Financial Lending: Models, Visualizations, and Summary-Explanations
- Learning Sparse Classifiers: Continuous and Mixed Integer Optimization Perspectives
- Interpretable Machine Learning Models for the Digital Clock Drawing Test
- The Backbone Method for Ultra-High Dimensional Sparse Machine Learning
- Interpretability with Accurate Small Models
- Learning Fair Rule Lists
- Fairness Warnings and Fair-MAML: Learning Fairly with Minimal Data
- Automation of Quantum Dot Measurement Analysis via Explainable Machine Learning
- Gaining Free or Low-Cost Transparency with Interpretable Partial Substitute
- Learning Hybrid Interpretable Models: Theory, Taxonomy, and Methods
- Concave Quadratic Cuts for Mixed-Integer Quadratic Problems
- Learning Channel Importance for High Content Imaging with Interpretable Deep Input Channel Mixing
- From Predictions to Decisions: Using Lookahead Regularization
- Shapley variable importance clouds for interpretable machine learning
- PreCog: Improving Crowdsourced Data Quality Before Acquisition
- Optimal Explanations of Linear Models
- Interactivity and Transparency in Medical Risk Assessment with Supersparse Linear Integer Models
- Heaviside Set Constrained Optimization: Optimality and Newton Method
- Learning Interpretable Concept-Based Models with Human Feedback
- CACTUS: Detecting and Resolving Conflicts in Objective Functions
- How to improve the interpretability of kernel learning
- Model-Agnostic Explanations using Minimal Forcing Subsets