collaborators

5 papers

cs.CY2025

First-Person Fairness in Chatbots

Tyna Eloundou, Alex Beutel, David G. Robinson +7

Evaluating chatbot fairness is crucial given their rapid proliferation, yet typical chatbot tasks (e.g., resume writing, entertainment) diverge from the institutional decision-maki…

cs.CL2025

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification

Yashwanth M., Vaibhav Singh, Ayush Maheshwari +2

We propose ARISE, a framework that iteratively induces rules and generates synthetic data for text classification. We combine synthetic data generation and automatic rule induction…

cs.LG2024

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Alex Beutel, Kai Xiao, Johannes Heidecke +1

Automated red teaming can discover rare model failures and generate challenging examples that can be used for training or evaluation. However, a core challenge in automated red tea…

cs.AI2024

Rule Based Rewards for Language Model Safety

Tong Mu, Alec Helyar, Johannes Heidecke +7

Reinforcement learning based fine-tuning of large language models (LLMs) on human preferences has been shown to enhance both their capabilities and safety behavior. However, in cas…

cs.CR2024

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Eric Wallace, Kai Xiao, Reimar Leike +3

Today's LLMs are susceptible to prompt injections, jailbreaks, and other attacks that allow adversaries to overwrite a model's original instructions with their own malicious prompt…