Learning Optimized Risk Scores
arXiv:1610.00168
Abstract
Risk scores are simple classification models that let users make quick risk predictions by adding and subtracting a few small numbers. These models are widely used in medicine and criminal justice, but are difficult to learn from data because they need to be calibrated, sparse, use small integer coefficients, and obey application-specific operational constraints. In this paper, we present a new machine learning approach to learn risk scores. We formulate the risk score problem as a mixed integer nonlinear program, and present a cutting plane algorithm for non-convex settings to efficiently recover its optimal solution. We improve our algorithm with specialized techniques to generate feasible solutions, narrow the optimality gap, and reduce data-related computation. Our approach can fit risk scores in a way that scales linearly in the number of samples, provides a certificate of optimality, and obeys real-world constraints without parameter tuning or post-processing. We benchmark the performance benefits of this approach through an extensive set of numerical experiments, comparing to risk scores built using heuristic approaches. We also discuss its practical benefits through a real-world application where we build a customized risk score for ICU seizure prediction in collaboration with the Massachusetts General Hospital.
Cited by in corpus (12)
- Deep learning for temporal data representation in electronic health records: A systematic review of challenges and methodologies
- Explainable Machine Learning for Public Policy: Use Cases, Gaps, and Research Directions
- AutoScore-Survival: Developing interpretable machine learning-based time-to-event scores with right-censored survival data
- TimberTrek: Exploring and Curating Sparse Decision Trees with Interactive Visualization
- A Holistic Approach to Interpretability in Financial Lending: Models, Visualizations, and Summary-Explanations
- Interpretable and Fair Boolean Rule Sets via Column Generation
- When Does Uncertainty Matter?: Understanding the Impact of Predictive Uncertainty in ML Assisted Decision Making
- Fairness Warnings and Fair-MAML: Learning Fairly with Minimal Data
- Fast and Interpretable Mortality Risk Scores for Critical Care Patients
- Connecting Interpretability and Robustness in Decision Trees through Separation
- Fair Decision Rules for Binary Classification
- Shapley variable importance clouds for interpretable machine learning