1 citations · 1 across the 8 of their papers we have counts for
10 papers
Debate Training Reduces Reward Hacking in RLAIF
Zachary Kenton, Lili Janzer, Rory Greig +8
We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking comp…
Human-AI Complementarity: A Goal for Amplified Oversight
Rishub Jain, Sophie Bridgers, Lili Janzer +3
Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes…
"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives
Ding Wang, Mark Díaz, Charvi Rastogi +10
Understanding what constitutes safety in AI-generated content is complex. While developers often rely on predefined taxonomies, real-world safety judgments also involve personal, s…
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
Charvi Rastogi, Tian Huey Teh, Pushkar Mishra +10
Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understand…
Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity
Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl +9
Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safe…
Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups
Charvi Rastogi, Tian Huey Teh, Pushkar Mishra +10
AI systems crucially rely on human ratings, but these ratings are often aggregated, obscuring the inherent diversity of perspectives in real-world phenomenon. This is particularly…