3 papers
cs.CL2026
Training Deliberative Monitors for Black-Box Scheming Detection
Aditya Sinha, Akshat Naik, Victor Gillioz +5
As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a central AI control problem. Existing…
cs.AI2026
Log analysis is necessary for credible evaluation of AI agents
Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8
Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated or deflated by shortcuts and…
cs.CL2025
Large Language Models Often Know When They Are Being Evaluated
Joe Needham, Giles Edkins, Govind Pimpale +2
If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior durin…