most citedSleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

39 citations · 40 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG2024

Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach

Tony T. Wang, John Hughes, Henry Sleight +7

Defending large language models against jailbreaks so that they never engage in a broadly-defined set of forbidden behaviors is an open problem. In this paper, we investigate the d…

cs.CL20241 cited

Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats

Jiaxin Wen, Vivek Hebbar, Caleb Larson +9

As large language models (LLMs) become increasingly capable, it is prudent to assess whether safety measures remain effective even if LLMs intentionally try to bypass them. Previou…

cs.CL2024

Rapid Response: Mitigating LLM Jailbreaks with a Few Examples

Alwin Peng, Julian Michael, Henry Sleight +2

As large language models (LLMs) grow more powerful, ensuring their safety against misuse becomes crucial. While researchers have focused on developing robust defenses, no method ha…

cs.CR202439 cited

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Evan Hubinger, Carson Denison, Jesse Mu +36

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when giv…

cs.AI2023

Understanding and Controlling a Maze-Solving Policy Network

Ulisse Mini, Peli Grietzer, Mrinank Sharma +3

To understand the goals and goal representations of AI systems, we carefully study a pretrained reinforcement learning policy that solves mazes by navigating to a range of target s…