activity
20172024
most citedFooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

168 citations · 302 across the 14 of their papers we have counts for

collaborators

19 papers

cs.LG2023

Consistent Explanations in the Face of Model Indeterminacy via Ensembling

Dan Ley, Leonard Tang, Matthew Nazari +3

This work addresses the challenge of providing consistent explanations for predictive models in the presence of model indeterminacy, which arises due to the existence of multiple (…

cs.LG20222 cited

Towards Robust Off-Policy Evaluation via Human Inputs

Harvineet Singh, Shalmali Joshi, Finale Doshi-Velez +1

Off-policy Evaluation (OPE) methods are crucial tools for evaluating policies in high-stakes domains such as healthcare, where direct deployment is often infeasible, unethical, or…

cs.LG20224 cited

Rethinking Stability for Attribution-based Explanations

Chirag Agarwal, Nari Johnson, Martin Pawelczyk +4

As attribution-based explanation methods are increasingly used to establish model trustworthiness in high-stakes situations, it is critical to ensure that these explanations are st…

cs.LG202228 cited

Rethinking Explainability as a Dialogue: A Practitioner's Perspective

Himabindu Lakkaraju, Dylan Slack, Yuxin Chen +2

As practitioners increasingly deploy machine learning models in critical domains such as health care, finance, and policy, it becomes vital to ensure that domain experts function e…

cs.LG20214 cited

Feature Attributions and Counterfactual Explanations Can Be Manipulated

Dylan Slack, Sophie Hilgard, Sameer Singh +1

As machine learning models are increasingly used in critical decision-making settings (e.g., healthcare, finance), there has been a growing emphasis on developing methods to explai…

cs.LG20216 cited

What will it take to generate fairness-preserving explanations?

Jessica Dai, Sohini Upadhyay, Stephen H. Bach +1

In situations where explanations of black-box models may be useful, the fairness of the black-box is also often a relevant concern. However, the link between the fairness of the bl…