3.2k citations · 4.2k across the 9 of their papers we have counts for
7 papers · 1 filter
Towards Automatic Concept-based Explanations
Amirata Ghorbani, James Wexler, James Zou +1
Interpretability has become an important topic of research as more machine learning (ML) models are deployed and widely used to make important decisions. Most of the current explan…
Proceedings of the 2018 ICML Workshop on Human Interpretability in Machine Learning (WHI 2018)
Been Kim, Kush R. Varshney, Adrian Weller
This is the Proceedings of the 2018 ICML Workshop on Human Interpretability in Machine Learning (WHI 2018), which was held in Stockholm, Sweden, July 14, 2018. Invited speakers wer…
To Trust Or Not To Trust A Classifier
Heinrich Jiang, Been Kim, Melody Y. Guan +1
Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has…
Human-in-the-Loop Interpretability Prior
Isaac Lage, Andrew Slavin Ross, Been Kim +2
We often desire our models to be interpretable as well as accurate. Prior work on optimizing models for interpretability has relied on easy-to-quantify proxies for interpretability…
The (Un)reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo +5
Saliency methods aim to explain the predictions of deep neural networks. These methods lack reliability when the explanation is sensitive to factors that do not contribute to the m…
Proceedings of the 2017 ICML Workshop on Human Interpretability in Machine Learning (WHI 2017)
Been Kim, Dmitry M. Malioutov, Kush R. Varshney +1
This is the Proceedings of the 2017 ICML Workshop on Human Interpretability in Machine Learning (WHI 2017), which was held in Sydney, Australia, August 10, 2017. Invited speakers w…