3.2k citations · 4.4k across the 16 of their papers we have counts for
6 papers · 2 filters
On Completeness-aware Concept-Based Explanations in Deep Neural Networks
Chih-Kuan Yeh, Been Kim, Sercan O. Arik +3
Human explanations of high-level decisions are often expressed in terms of key concepts the decisions are based on. In this paper, we study such concept-based explainability for De…
Towards Realistic Individual Recourse and Actionable Explanations in Black-Box Decision Making Systems
Shalmali Joshi, Oluwasanmi Koyejo, Warut Vijitbenjaronk +2
Machine learning based decision making systems are increasingly affecting humans. An individual can suffer an undesirable outcome under such decision making systems (e.g. denied cr…
Benchmarking Attribution Methods with Relative Feature Importance
Mengjiao Yang, Been Kim
Interpretability is an important area of research for safe deployment of machine learning systems. One particular type of interpretability method attributes model decisions to inpu…
Explaining Classifiers with Causal Concept Effect (CaCE)
Yash Goyal, Amir Feder, Uri Shalit +1
How can we understand classification decisions made by deep neural networks? Many existing explainability methods rely solely on correlations and fail to account for confounding, w…
Visualizing and Measuring the Geometry of BERT
Andy Coenen, Emily Reif, Ann Yuan +4
Transformer architectures show significant promise for natural language processing. Given that a single pretrained model can be fine-tuned to perform well on many different tasks,…
Neural Networks Trained on Natural Scenes Exhibit Gestalt Closure
Been Kim, Emily Reif, Martin Wattenberg +2
The Gestalt laws of perceptual organization, which describe how visual elements in an image are grouped and interpreted, have traditionally been thought of as innate despite their…