5 papers
First-Person Fairness in Chatbots
Tyna Eloundou, Alex Beutel, David G. Robinson +7
Evaluating chatbot fairness is crucial given their rapid proliferation, yet typical chatbot tasks (e.g., resume writing, entertainment) diverge from the institutional decision-maki…
ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
Yashwanth M., Vaibhav Singh, Ayush Maheshwari +2
We propose ARISE, a framework that iteratively induces rules and generates synthetic data for text classification. We combine synthetic data generation and automatic rule induction…
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Alex Beutel, Kai Xiao, Johannes Heidecke +1
Automated red teaming can discover rare model failures and generate challenging examples that can be used for training or evaluation. However, a core challenge in automated red tea…
Rule Based Rewards for Language Model Safety
Tong Mu, Alec Helyar, Johannes Heidecke +7
Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cas…
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Eric Wallace, Kai Xiao, Reimar Leike +3
Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompt…