2 papers
cs.AI2026
One Probe Won't Catch Them All: Towards Targeted Deception Detection
Vikram Natarajan, Devina Jain, Shivam Arora +2
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier trained on a contrastive instruction pair…
cs.AI2025
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…