Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Scaling Trends for Lie Detector Oversight in Preference Learning
Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day +2
Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detec…
cs.AI2025
Neural Interactive Proofs
Lewis Hammond, Sam Adam-Day
We consider the problem of how a trusted, but computationally bounded agent (a 'verifier') can learn to interact with one or more powerful but untrusted agents ('provers') in order…