430 citations · 912 across the 21 of their papers we have counts for
21 papers · 1 filter
Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift
Benjamin Eyre, Elliot Creager, David Madras +2
Designing deep neural network classifiers that perform robustly on distributions differing from the available training data is an active area of machine learning research. However,…
Improving Few-shot Generalization of Safety Classifiers via Data Augmented Parameter-Efficient Fine-Tuning
Ananth Balashankar, Xiao Ma, Aradhana Sinha +4
As large language models (LLMs) are widely adopted, new safety issues and policies emerge, to which existing safety classifiers do not generalize well. If we have only observed a f…
Controlled Decoding from Language Models
Sidharth Mudgal, Jong Lee, Harish Ganapathy +10
KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective a…
Break it, Imitate it, Fix it: Robustness by Generating Human-Like Attacks
Aradhana Sinha, Ananth Balashankar, Ahmad Beirami +3
Real-world natural language processing systems need to be robust to human adversaries. Collecting examples of human adversaries for training is an effective but expensive solution.…
Towards A Scalable Solution for Improving Multi-Group Fairness in Compositional Classification
James Atwood, Tina Tian, Ben Packer +5
Despite the rich literature on machine learning fairness, relatively little attention has been paid to remediating complex systems, where the final prediction is the combination of…
Striving for data-model efficiency: Identifying data externalities on group performance
Esther Rolf, Ben Packer, Alex Beutel +1
Building trustworthy, effective, and responsible machine learning systems hinges on understanding how differences in training data and modeling decisions interact to impact predict…