activity
20172022
most citedTowards A Rigorous Science of Interpretable Machine Learning

3.2k citations · 4.2k across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

12 papers · 1 filter

cs.LG202225 cited

Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation

Julius Adebayo, Michael Muelly, Hal Abelson +1

We investigate whether three types of post hoc model explanations--feature attribution, concept activation, and training point ranking--are effective for detecting a model's relian…

cs.LG202212 cited

Human-Centered Concept Explanations for Neural Networks

Chih-Kuan Yeh, Been Kim, Pradeep Ravikumar

Understanding complex machine learning models such as deep neural networks with explanations is crucial in various applications. Many explanations stem from the model perspective,…

cs.LG2020

Concept Bottleneck Models

Pang Wei Koh, Thao Nguyen, Yew Siang Tang +4

We seek to learn models that we can interact with using high-level concepts: if the model did not think there was a bone spur in the x-ray, would it still predict severe arthritis?…

cs.LG201997 cited

Towards Realistic Individual Recourse and Actionable Explanations in Black-Box Decision Making Systems

Shalmali Joshi, Oluwasanmi Koyejo, Warut Vijitbenjaronk +2

Machine learning based decision making systems are increasingly affecting humans. An individual can suffer an undesirable outcome under such decision making systems (e.g. denied cr…

cs.LG2019

Benchmarking Attribution Methods with Relative Feature Importance

Mengjiao Yang, Been Kim

Interpretability is an important area of research for safe deployment of machine learning systems. One particular type of interpretability method attributes model decisions to inpu…

cs.LG2019

Explaining Classifiers with Causal Concept Effect (CaCE)

Yash Goyal, Amir Feder, Uri Shalit +1

How can we understand classification decisions made by deep neural networks? Many existing explainability methods rely solely on correlations and fail to account for confounding, w…