5 papers
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Alan Cooney, David Africa, Geoffrey Irving
Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires test…
Behavioural Analysis of Alignment Faking
Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1
Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understandi…
Practical challenges of control monitoring in frontier AI deployments
David Lindner, Charlie Griffin, Tomek Korbak +4
Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified…
Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
Asa Cooper Stickland, Jan Michelfeit, Arathi Mani +6
LLM-based software engineering agents are increasingly used in real-world development tasks, often with access to sensitive data or security-critical codebases. Such agents could i…
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
Sid Black, Asa Cooper Stickland, Jake Pencharz +7
Uncontrollable autonomous replication of language model agents poses a critical safety risk. To better understand this risk, we introduce RepliBench, a suite of evaluations designe…