Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Scaling Trends for Lie Detector Oversight in Preference Learning
Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day +2
Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detec…
cs.AI2025
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…