208 citations · 219 across the 6 of their papers we have counts for
7 papers
Rethinking Stability for Attribution-based Explanations
Chirag Agarwal, Nari Johnson, Martin Pawelczyk +4
As attribution-based explanation methods are increasingly used to establish model trustworthiness in high-stakes situations, it is critical to ensure that these explanations are st…
Does Robustness Improve Fairness? Approaching Fairness with Word Substitution Robustness Methods for Text Classification
Yada Pruksachatkun, Satyapriya Krishna, Jwala Dhamala +2
Existing bias mitigation methods to reduce disparities in model outcomes across cohorts have focused on data augmentation, debiasing model embeddings, or adding fairness-based opti…
Grounding Complex Navigational Instructions Using Scene Graphs
Michiel de Jong, Satyapriya Krishna, Anuva Agarwal
Training a reinforcement learning agent to carry out natural language instructions is limited by the available supervision, i.e. knowing when the instruction has been carried out.…
ADePT: Auto-encoder based Differentially Private Text Transformation
Satyapriya Krishna, Rahul Gupta, Christophe Dupuy
Privacy is an important concern when building statistical models on data containing personal information. Differential privacy offers a strong definition of privacy and can be used…
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Jwala Dhamala, Tony Sun, Varun Kumar +4
Recent advances in deep learning techniques have enabled machines to generate cohesive open-ended text when prompted with a sequence of words as context. While these models now emp…
Towards classification parity across cohorts
Aarsh Patel, Rahul Gupta, Mukund Harakere +3
Recently, there has been a lot of interest in ensuring algorithmic fairness in machine learning where the central question is how to prevent sensitive information (e.g. knowledge a…