Learning Interpretable Models with Causal Guarantees
arXiv:1901.08576
Abstract
Machine learning has shown much promise in helping improve the quality of medical, legal, and financial decision-making. In these applications, machine learning models must satisfy two important criteria: (i) they must be causal, since the goal is typically to predict individual treatment effects, and (ii) they must be interpretable, so that human decision makers can validate and trust the model predictions. There has recently been much progress along each direction independently, yet the state-of-the-art approaches are fundamentally incompatible. We propose a framework for learning interpretable models from observational data that can be used to predict individual treatment effects (ITEs). In particular, our framework converts any supervised learning algorithm into an algorithm for estimating ITEs. Furthermore, we prove an error bound on the treatment effects predicted by our model. Finally, in an experiment on real-world data, we show that the models trained using our framework significantly outperform a number of baselines.
References in corpus (6)
- Distilling the Knowledge in a Neural Network
- Towards A Rigorous Science of Interpretable Machine Learning
- Equality of Opportunity in Supervised Learning
- Distilling a Neural Network Into a Soft Decision Tree
- Verifiable Reinforcement Learning via Policy Extraction
- HOUDINI: Lifelong Learning as Program Synthesis
Cited by in corpus (7)
- Evaluating Explanation Without Ground Truth in Interpretable Machine Learning
- Artificial Intelligence in Materials Science and Engineering: Current Landscape, Key Challenges, and Future Trajectorie
- Generative causal explanations of black-box classifiers
- Causal Interpretability for Machine Learning -- Problems, Methods and Evaluation
- Robust and Stable Black Box Explanations
- Information-theoretic Evolution of Model Agnostic Global Explanations
- Causal Explanations of Image Misclassifications