collaborators

7 papers

cs.AI2026

Scaling Trends for Lie Detector Oversight in Preference Learning

Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day +2

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detec…

cs.LG2026

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave +1

Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to ev…

cs.AI2025

Auditing Games for Sandbagging

Jordan Taylor, Sid Black, Dillon Bowen +10

Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…

cs.LG2025

Preference Learning with Lie Detectors can Induce Honesty or Evasion

Chris Cundy, Adam Gleave

As AI systems become more capable, deceptive behaviors can undermine evaluation and mislead users at deployment. Recent work has shown that lie detectors can accurately classify de…

cs.CY2025

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

Ann-Kathrin Dombrowski, Dillon Bowen, Adam Gleave +1

Open-weight large language models (LLMs) unlock huge benefits in innovation, personalization, privacy, and democratization. However, their core advantage - modifiability - opens th…

cs.LG2025

Planning in a recurrent neural network that plays Sokoban

Mohammad Taufeeque, Philip Quirke, Maximilian Li +4

Planning is essential for solving complex tasks, yet the internal mechanisms underlying planning in neural networks remain poorly understood. Building on prior work, we analyze a r…