activity
20192021
most citedFooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

168 citations · 177 across the 5 of their papers we have counts for

collaborators

9 papers

cs.LG20214 cited

Feature Attributions and Counterfactual Explanations Can Be Manipulated

Dylan Slack, Sophie Hilgard, Sameer Singh +1

As machine learning models are increasingly used in critical decision-making settings (e.g., healthcare, finance), there has been a growing emphasis on developing methods to explai…

cs.CL2021

On the Lack of Robust Interpretability of Neural Text Classifiers

Muhammad Bilal Zafar, Michele Donini, Dylan Slack +3

With the ever-increasing complexity of neural language models, practitioners have turned to methods for understanding the predictions of these models. One of the most well-adopted…

cs.LG20213 cited

Counterfactual Explanations Can Be Manipulated

Dylan Slack, Sophie Hilgard, Himabindu Lakkaraju +1

Counterfactual explanations are emerging as an attractive option for providing recourse to individuals adversely impacted by algorithmic decisions. As they are deployed in critical…

cs.LG20211 cited

Defuse: Harnessing Unrestricted Adversarial Examples for Debugging Models Beyond Test Accuracy

Dylan Slack, Nathalie Rauschmayr, Krishnaram Kenthapadi

We typically compute aggregate statistics on held-out test data to assess the generalization of machine learning models. However, statistics on test data often overstate model gene…

cs.LG2020

Differentially Private Language Models Benefit from Public Pre-training

Gavin Kerrigan, Dylan Slack, Jens Tuyls

Language modeling is a keystone task in natural language processing. When training a language model on sensitive information, differential privacy (DP) allows us to quantify the de…

cs.LG20191 cited

Fair Meta-Learning: Learning How to Learn Fairly

Dylan Slack, Sorelle Friedler, Emile Givental

Data sets for fairness relevant tasks can lack examples or be biased according to a specific label in a sensitive attribute. We demonstrate the usefulness of weight based meta-lear…