39 citations · 41 across the 3 of their papers we have counts for
3 papers
cs.CR2024★ 39 cited
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Evan Hubinger, Carson Denison, Jesse Mu +36
Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when giv…
cs.LG2023★ 2 cited
Preventing Language Models From Hiding Their Reasoning
Fabien Roger, Ryan Greenblatt
Large language models (LLMs) often benefit from intermediate steps of reasoning to generate answers to complex problems. When these intermediate steps of reasoning are used to moni…
cs.LG2023
Benchmarks for Detecting Measurement Tampering
Fabien Roger, Ryan Greenblatt, Max Nadeau +2
When training powerful AI systems to perform complex tasks, it may be challenging to provide training signals which are robust to optimization. One concern is \textit{measurement t…