papers

Publications (18)

cs.LG2023

TalkToModel: Explaining Machine Learning Models with Interactive Natural Language Conversations

Dylan Slack, Satyapriya Krishna, Himabindu Lakkaraju +1

Machine Learning (ML) models are increasingly used to make critical decisions in real-world applications, yet they have become more complex, making them harder to understand. To th…

cs.CL2024

Learning Goal-Conditioned Representations for Language Reward Models

Vaskar Nath, Dylan Slack, Jeff Da +4

Techniques that learn improved representations via offline data or self-supervised objectives have shown impressive results in traditional reinforcement learning (RL). Nevertheless…

cs.LG2021

Counterfactual Explanations Can Be Manipulated

Dylan Slack, Sophie Hilgard, Himabindu Lakkaraju +1

Counterfactual explanations are emerging as an attractive option for providing recourse to individuals adversely impacted by algorithmic decisions. As they are deployed in critical…

cs.AI2025

Towards physician-centered oversight of conversational diagnostic AI

Elahe Vedadi, David Barrett, Natalie Harris +32

Recent work has demonstrated the promise of conversational AI systems for diagnostic dialogue. However, real-world assurance of patient safety means that providing individual diagn…

cs.LG2019

Fairness Warnings and Fair-MAML: Learning Fairly with Minimal Data

Dylan Slack, Sorelle Friedler, Emile Givental

Motivated by concerns surrounding the fairness effects of sharing and transferring fair machine learning tools, we propose two algorithms: Fairness Warnings and Fair-MAML. The firs…

cs.LG2021

Defuse: Harnessing Unrestricted Adversarial Examples for Debugging Models Beyond Test Accuracy

Dylan Slack, Nathalie Rauschmayr, Krishnaram Kenthapadi

We typically compute aggregate statistics on held-out test data to assess the generalization of machine learning models. However, statistics on test data often overstate model gene…

cs.LG2020

Differentially Private Language Models Benefit from Public Pre-training

Gavin Kerrigan, Dylan Slack, Jens Tuyls

Language modeling is a keystone task in natural language processing. When training a language model on sensitive information, differential privacy (DP) allows us to quantify the de…

cs.LG2021

Reliable Post hoc Explanations: Modeling Uncertainty in Explainability

Dylan Slack, Sophie Hilgard, Sameer Singh +1

As black box explanations are increasingly being employed to establish model credibility in high-stakes settings, it is important to ensure that these explanations are accurate and…

cs.LG2019

Fair Meta-Learning: Learning How to Learn Fairly

Dylan Slack, Sorelle Friedler, Emile Givental

Data sets for fairness relevant tasks can lack examples or be biased according to a specific label in a sensitive attribute. We demonstrate the usefulness of weight based meta-lear…

cs.LG2020

Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods

Dylan Slack, Sophie Hilgard, Emily Jia +2

As machine learning black boxes are increasingly being deployed in domains such as healthcare and criminal justice, there is growing emphasis on building tools and techniques for e…

cs.LG2019

Assessing the Local Interpretability of Machine Learning Models

Dylan Slack, Sorelle A. Friedler, Carlos Scheidegger +1

The increasing adoption of machine learning tools has led to calls for accountability via model interpretability. But what does it mean for a machine learning model to be interpret…

cs.CL2021

On the Lack of Robust Interpretability of Neural Text Classifiers

Muhammad Bilal Zafar, Michele Donini, Dylan Slack +3

With the ever-increasing complexity of neural language models, practitioners have turned to methods for understanding the predictions of these models. One of the most well-adopted…

cs.LG2021

Feature Attributions and Counterfactual Explanations Can Be Manipulated

Dylan Slack, Sophie Hilgard, Sameer Singh +1

As machine learning models are increasingly used in critical decision-making settings (e.g., healthcare, finance), there has been a growing emphasis on developing methods to explai…

cs.CL2024

A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Hugh Zhang, Jeff Da, Dean Lee +12

Large language models (LLMs) have achieved impressive success on many benchmarks for mathematical reasoning. However, there is growing concern that some of this performance actuall…

cs.LG2022

SAFER: Data-Efficient and Safe Reinforcement Learning via Skill Acquisition

Dylan Slack, Yinlam Chow, Bo Dai +1

Methods that extract policy primitives from offline demonstrations using deep generative models have shown promise at accelerating reinforcement learning(RL) for new tasks. Intuiti…

cs.CL2023

Post Hoc Explanations of Language Models Can Improve Language Models

Satyapriya Krishna, Jiaqi Ma, Dylan Slack +3

Large Language Models (LLMs) have demonstrated remarkable capabilities in performing complex tasks. Moreover, recent research has shown that incorporating human-annotated rationale…

cs.LG2023

TABLET: Learning From Instructions For Tabular Data

Dylan Slack, Sameer Singh

Acquiring high-quality data is often a significant challenge in training machine learning (ML) models for tabular prediction, particularly in privacy-sensitive and costly domains l…

cs.LG2022

Rethinking Explainability as a Dialogue: A Practitioner's Perspective

Himabindu Lakkaraju, Dylan Slack, Yuxin Chen +2

As practitioners increasingly deploy machine learning models in critical domains such as health care, finance, and policy, it becomes vital to ensure that domain experts function e…