activity
20232026
most citedAlignment faking in large language models

24 citations · 29 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG20251 cited

Why Do Some Language Models Fake Alignment While Others Don't?

Abhay Sheshadri, John Hughes, Julian Michael +4

Alignment faking in large language models presented a demonstration of Claude 3 Opus and Claude 3.5 Sonnet selectively complying with a helpful-only training objective to prevent m…

cs.LG2024

Stress-Testing Capability Elicitation With Password-Locked Models

Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov +1

To determine the safety of large language models (LLMs), AI developers must be able to assess their dangerous capabilities. But simple prompting strategies often fail to elicit an…

cs.LG2023

AI Control: Improving Safety Despite Intentional Subversion

Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan +1

As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researcher…

cs.LG20232 cited

Preventing Language Models From Hiding Their Reasoning

Fabien Roger, Ryan Greenblatt

Large language models (LLMs) often benefit from intermediate steps of reasoning to generate answers to complex problems. When these intermediate steps of reasoning are used to moni…

cs.LG2023

Benchmarks for Detecting Measurement Tampering

Fabien Roger, Ryan Greenblatt, Max Nadeau +2

When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization. One concern is \textit{measurement t…