Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Training Deliberative Monitors for Black-Box Scheming Detection
Aditya Sinha, Akshat Naik, Victor Gillioz +5
As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a central AI control problem. Existing…
cs.CL2025
Large Language Models Often Know When They Are Being Evaluated
Joe Needham, Giles Edkins, Govind Pimpale +2
If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior durin…