3 papers
cs.AI2026
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Alan Cooney, David Africa, Geoffrey Irving
Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model behaviour, but evaluating them requires test…
cs.AI2026
Debate is efficient with your time
Jonah Brown-Cohen, Geoffrey Irving, Simon C. Marshall +3
AI safety via debate uses two competing models to help a human judge verify complex computational tasks. Previous work has established what problems debate can solve in principle,…
cs.AI2025
Avoiding Obfuscation with Prover-Estimator Debate
Jonah Brown-Cohen, Geoffrey Irving, Georgios Piliouras +3
Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this pr…