activity
20172024
most citedUnderspecification Presents Challenges for Credibility in Modern Machine Learning

430 citations · 912 across the 21 of their papers we have counts for

collaborators
Showing cs.LGShow all

21 papers · 1 filter

cs.LG2023

Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift

Benjamin Eyre, Elliot Creager, David Madras +2

Designing deep neural network classifiers that perform robustly on distributions differing from the available training data is an active area of machine learning research. However,…

cs.LG2023

Improving Few-shot Generalization of Safety Classifiers via Data Augmented Parameter-Efficient Fine-Tuning

Ananth Balashankar, Xiao Ma, Aradhana Sinha +4

As large language models (LLMs) are widely adopted, new safety issues and policies emerge, to which existing safety classifiers do not generalize well. If we have only observed a f…

cs.LG2023

Controlled Decoding from Language Models

Sidharth Mudgal, Jong Lee, Harish Ganapathy +10

KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective a…

cs.LG2023

Break it, Imitate it, Fix it: Robustness by Generating Human-Like Attacks

Aradhana Sinha, Ananth Balashankar, Ahmad Beirami +3

Real-world natural language processing systems need to be robust to human adversaries. Collecting examples of human adversaries for training is an effective but expensive solution.…

cs.LG2023

Towards A Scalable Solution for Improving Multi-Group Fairness in Compositional Classification

James Atwood, Tina Tian, Ben Packer +5

Despite the rich literature on machine learning fairness, relatively little attention has been paid to remediating complex systems, where the final prediction is the combination of…

cs.LG2022

Striving for data-model efficiency: Identifying data externalities on group performance

Esther Rolf, Ben Packer, Alex Beutel +1

Building trustworthy, effective, and responsible machine learning systems hinges on understanding how differences in training data and modeling decisions interact to impact predict…