4 papers · 1 filter
One Probe Won't Catch Them All: Towards Targeted Deception Detection
Vikram Natarajan, Devina Jain, Shivam Arora +2
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair…
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
Satvik Golechha, Adrià Garriga-Alonso
Prior studies on deception in language-based AI agents typically assess whether the agent produces a false statement about a topic, or makes a binary choice prompted by a goal, rat…
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…
Auditing language models for hidden objectives
Samuel Marks, Johannes Treutlein, Trenton Bricken +32
We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…