11 citations · 26 across the 11 of their papers we have counts for
11 papers
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Eric Wallace, Kai Xiao, Reimar Leike +3
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompt…
Generalized People Diversity: Learning a Human Perception-Aligned Diversity Representation for People Images
Hansa Srinivasan, Candice Schumann, Aradhana Sinha +5
Capturing the diversity of people in images is challenging: recent literature tends to focus on diversifying one or two attributes, requiring expensive attribute labels or building…
Improving Few-shot Generalization of Safety Classifiers via Data Augmented Parameter-Efficient Fine-Tuning
Ananth Balashankar, Xiao Ma, Aradhana Sinha +4
As large language models (LLMs) are widely adopted, new safety issues and policies emerge, to which existing safety classifiers do not generalize well. If we have only observed a f…
Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting
Preethi Lahoti, Nicholas Blumm, Xiao Ma +8
A crucial challenge for generative large language models (LLMs) is diversity: when a user's prompt is under-specified, models may follow implicit assumptions while generating a res…
Learning from Negative User Feedback and Measuring Responsiveness for Sequential Recommenders
Yueqi Wang, Yoni Halpern, Shuo Chang +9
Sequential recommenders have been widely used in industry due to their strength in modeling user preferences. While these models excel at learning a user's positive interests, less…
Towards A Scalable Solution for Improving Multi-Group Fairness in Compositional Classification
James Atwood, Tina Tian, Ben Packer +5
Despite the rich literature on machine learning fairness, relatively little attention has been paid to remediating complex systems, where the final prediction is the combination of…