most citedSleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

39 citations · 104 across the 6 of their papers we have counts for

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI20253 cited

Auditing language models for hidden objectives

Samuel Marks, Johannes Treutlein, Trenton Bricken +32

We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…

cs.AI202424 cited

Alignment faking in large language models

Ryan Greenblatt, Carson Denison, Benjamin Wright +17

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…

cs.AI20249 cited

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Carson Denison, Monte MacDiarmid, Fazl Barez +11

In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals. Specification gaming c…

cs.AI202329 cited

Measuring Faithfulness in Chain-of-Thought Reasoning

Tamera Lanham, Anna Chen, Ansh Radhakrishnan +27

Large language models (LLMs) perform better when they produce step-by-step, "Chain-of-Thought" (CoT) reasoning before answering a question, but it is unclear if the stated reasonin…

cs.AI2023

Conditioning Predictive Models: Risks and Strategies

Evan Hubinger, Adam Jermyn, Johannes Treutlein +2

Our intention is to provide a definitive reference on what it would take to safely make use of generative/predictive models in the absence of a solution to the Eliciting Latent Kno…