2 papers
cs.AI2026
Evaluating and Understanding Scheming Propensity in LLM Agents
Mia Hopman, Jannes Elstner, Maria Avramidou +2
As frontier language models are increasingly deployed as autonomous agents pursuing complex, long-term objectives, there is increased risk of scheming: agents covertly pursuing mis…
cs.CY2025
Combining Cost-Constrained Runtime Monitors for AI Safety
Tim Tian Hua, James Baskerville, Henri Lemoine +3
Monitoring AIs at runtime can help us detect and stop harmful actions. In this paper, we study how to efficiently combine multiple runtime monitors into a single monitoring protoco…