activity
20242026
most citedConstitutional Black-Box Monitoring for Scheming in LLM Agents

1 citations · 1 across the 1 of their papers we have counts for

collaborators

9 papers

cs.CL20261 cited

Constitutional Black-Box Monitoring for Scheming in LLM Agents

Simon Storf, Rich Barton-Cooper, James Peters-Gill +1

Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms. A central challenge is detecting scheming, where agents covertly…

cs.AI2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…

cs.AI2025

Stress Testing Deliberative Alignment for Anti-Scheming Training

Bronson Schoen, Evgenia Nitishinskaya, Mikita Balesni +16

Highly capable AI systems could secretly pursue misaligned goals -- what we call "scheming". Because a scheming AI would deliberately try to hide its misaligned goals and actions,…

cs.CY2025

AI Behind Closed Doors: a Primer on The Governance of Internal Deployment

Charlotte Stix, Matteo Pistillo, Girish Sastry +6

The most advanced future AI systems will first be deployed inside the frontier AI companies developing them. According to these companies and independent experts, AI systems may re…

cs.CL2025

Forecasting Frontier Language Model Agent Capabilities

Govind Pimpale, Axel Højmark, Jérémy Scheurer +1

As Language Models (LMs) increasingly operate as autonomous agents, accurately forecasting their capabilities becomes crucial for societal preparedness. We evaluate six forecasting…

cs.LG2025

Detecting Strategic Deception Using Linear Probes

Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim +1

AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs…