Publications (15)
Debugging Tests for Model Explanations
Julius Adebayo, Michael Muelly, Ilaria Liccardi +1
We investigate whether post-hoc model explanations are effective for diagnosing model errors--model debugging. In response to the challenge of explaining a model's prediction, a va…
Iterative Orthogonal Feature Projection for Diagnosing Bias in Black-Box Models
Julius Adebayo, Lalana Kagal
Predictive models are increasingly deployed for the purpose of determining access to services such as credit, insurance, and employment. Despite potential gains in productivity and…
Generative Models, Humans, Predictive Models: Who Is Worse at High-Stakes Decision Making?
Keri Mallari, Julius Adebayo, Kori Inkpen +3
Despite strong advisory against it, large generative models (LMs) are already being used for decision making tasks that were previously done by predictive models or humans. We put…
Explaining Explanations to Society
Leilani H. Gilpin, Cecilia Testart, Nathaniel Fruchter +1
There is a disconnect between explanatory artificial intelligence (XAI) methods and the types of explanations that are useful for and demanded by society (policy makers, government…
Concept Bottleneck Language Models For protein design
Aya Abdelsalam Ismail, Tuomas Oikarinen, Amy Wang +8
We introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our arc…
Error Discovery by Clustering Influence Embeddings
Fulton Wang, Julius Adebayo, Sarah Tan +2
We present a method for identifying groups of test examples -- slices -- on which a model under-performs, a task now known as slice discovery. We formalize coherence -- a requireme…