168 citations · 177 across the 5 of their papers we have counts for
9 papers
Feature Attributions and Counterfactual Explanations Can Be Manipulated
Dylan Slack, Sophie Hilgard, Sameer Singh +1
As machine learning models are increasingly used in critical decision-making settings (e.g., healthcare, finance), there has been a growing emphasis on developing methods to explai…
On the Lack of Robust Interpretability of Neural Text Classifiers
Muhammad Bilal Zafar, Michele Donini, Dylan Slack +3
With the ever-increasing complexity of neural language models, practitioners have turned to methods for understanding the predictions of these models. One of the most well-adopted…
Counterfactual Explanations Can Be Manipulated
Dylan Slack, Sophie Hilgard, Himabindu Lakkaraju +1
Counterfactual explanations are emerging as an attractive option for providing recourse to individuals adversely impacted by algorithmic decisions. As they are deployed in critical…
Defuse: Harnessing Unrestricted Adversarial Examples for Debugging Models Beyond Test Accuracy
Dylan Slack, Nathalie Rauschmayr, Krishnaram Kenthapadi
We typically compute aggregate statistics on held-out test data to assess the generalization of machine learning models. However, statistics on test data often overstate model gene…
Differentially Private Language Models Benefit from Public Pre-training
Gavin Kerrigan, Dylan Slack, Jens Tuyls
Language modeling is a keystone task in natural language processing. When training a language model on sensitive information, differential privacy (DP) allows us to quantify the de…
Fair Meta-Learning: Learning How to Learn Fairly
Dylan Slack, Sorelle Friedler, Emile Givental
Data sets for fairness relevant tasks can lack examples or be biased according to a specific label in a sensitive attribute. We demonstrate the usefulness of weight based meta-lear…