299 citations · 412 across the 19 of their papers we have counts for
14 papers · 1 filter
Improving Few-shot Generalization of Safety Classifiers via Data Augmented Parameter-Efficient Fine-Tuning
Ananth Balashankar, Xiao Ma, Aradhana Sinha +4
As large language models (LLMs) are widely adopted, new safety issues and policies emerge, to which existing safety classifiers do not generalize well. If we have only observed a f…
Controlled Decoding from Language Models
Sidharth Mudgal, Jong Lee, Harish Ganapathy +10
KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective a…
Break it, Imitate it, Fix it: Robustness by Generating Human-Like Attacks
Aradhana Sinha, Ananth Balashankar, Ahmad Beirami +3
Real-world natural language processing systems need to be robust to human adversaries. Collecting examples of human adversaries for training is an effective but expensive solution.…
Towards A Scalable Solution for Improving Multi-Group Fairness in Compositional Classification
James Atwood, Tina Tian, Ben Packer +5
Despite the rich literature on machine learning fairness, relatively little attention has been paid to remediating complex systems, where the final prediction is the combination of…
A Human-ML Collaboration Framework for Improving Video Content Reviews
Meghana Deodhar, Xiao Ma, Yixin Cai +3
We deal with the problem of localized in-video taxonomic human annotation in the video content moderation domain, where the goal is to identify video segments that violate granular…
Understanding and Improving Fairness-Accuracy Trade-offs in Multi-Task Learning
Yuyan Wang, Xuezhi Wang, Alex Beutel +3
As multi-task models gain popularity in a wider range of machine learning applications, it is becoming increasingly important for practitioners to understand the fairness implicati…