391 citations · 858 across the 14 of their papers we have counts for
9 papers · 1 filter
Jailbreak Distillation: Renewable Safety Benchmarking
Jingyu Zhang, Ahmed Elgohary, Xiawei Wang +5
Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a no…
Specific versus General Principles for Constitutional AI
Sandipan Kundu, Yuntao Bai, Saurav Kadavath +33
Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-pre…
Towards Understanding Sycophancy in Language Models
Mrinank Sharma, Meg Tong, Tomasz Korbak +16
Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known a…
Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
Ansh Radhakrishnan, Karina Nguyen, Anna Chen +21
As large language models (LLMs) perform more difficult tasks, it becomes harder to verify the correctness and safety of their behavior. One approach to help with this issue is to p…
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Esin Durmus, Karina Nguyen, Thomas I. Liao +15
Large language models (LLMs) may not equitably represent diverse global perspectives on societal issues. In this paper, we develop a quantitative framework to evaluate whose opinio…
The Capacity for Moral Self-Correction in Large Language Models
Deep Ganguli, Amanda Askell, Nicholas Schiefer +46
We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmf…