activity
20192022
most citedFooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

168 citations · 205 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG202228 cited

Rethinking Explainability as a Dialogue: A Practitioner's Perspective

Himabindu Lakkaraju, Dylan Slack, Yuxin Chen +2

As practitioners increasingly deploy machine learning models in critical domains such as health care, finance, and policy, it becomes vital to ensure that domain experts function e…

cs.LG20214 cited

Feature Attributions and Counterfactual Explanations Can Be Manipulated

Dylan Slack, Sophie Hilgard, Sameer Singh +1

As machine learning models are increasingly used in critical decision-making settings (e.g., healthcare, finance), there has been a growing emphasis on developing methods to explai…

cs.LG20213 cited

Counterfactual Explanations Can Be Manipulated

Dylan Slack, Sophie Hilgard, Himabindu Lakkaraju +1

Counterfactual explanations are emerging as an attractive option for providing recourse to individuals adversely impacted by algorithmic decisions. As they are deployed in critical…

cs.LG20211 cited

Defuse: Harnessing Unrestricted Adversarial Examples for Debugging Models Beyond Test Accuracy

Dylan Slack, Nathalie Rauschmayr, Krishnaram Kenthapadi

We typically compute aggregate statistics on held-out test data to assess the generalization of machine learning models. However, statistics on test data often overstate model gene…

cs.LG2020

Differentially Private Language Models Benefit from Public Pre-training

Gavin Kerrigan, Dylan Slack, Jens Tuyls

Language modeling is a keystone task in natural language processing. When training a language model on sensitive information, differential privacy (DP) allows us to quantify the de…

cs.LG20191 cited

Fair Meta-Learning: Learning How to Learn Fairly

Dylan Slack, Sorelle Friedler, Emile Givental

Data sets for fairness relevant tasks can lack examples or be biased according to a specific label in a sensitive attribute. We demonstrate the usefulness of weight based meta-lear…