3.2k citations · 4.2k across the 9 of their papers we have counts for
12 papers · 1 filter
Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation
Julius Adebayo, Michael Muelly, Hal Abelson +1
We investigate whether three types of post hoc model explanations--feature attribution, concept activation, and training point ranking--are effective for detecting a model's relian…
Human-Centered Concept Explanations for Neural Networks
Chih-Kuan Yeh, Been Kim, Pradeep Ravikumar
Understanding complex machine learning models such as deep neural networks with explanations is crucial in various applications. Many explanations stem from the model perspective,…
Concept Bottleneck Models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang +4
We seek to learn models that we can interact with using high-level concepts: if the model did not think there was a bone spur in the x-ray, would it still predict severe arthritis?…
Towards Realistic Individual Recourse and Actionable Explanations in Black-Box Decision Making Systems
Shalmali Joshi, Oluwasanmi Koyejo, Warut Vijitbenjaronk +2
Machine learning based decision making systems are increasingly affecting humans. An individual can suffer an undesirable outcome under such decision making systems (e.g. denied cr…
Benchmarking Attribution Methods with Relative Feature Importance
Mengjiao Yang, Been Kim
Interpretability is an important area of research for safe deployment of machine learning systems. One particular type of interpretability method attributes model decisions to inpu…
Explaining Classifiers with Causal Concept Effect (CaCE)
Yash Goyal, Amir Feder, Uri Shalit +1
How can we understand classification decisions made by deep neural networks? Many existing explainability methods rely solely on correlations and fail to account for confounding, w…