32 citations · 80 across the 12 of their papers we have counts for
12 papers
Stress Testing Deliberative Alignment for Anti-Scheming Training
Bronson Schoen, Evgenia Nitishinskaya, Mikita Balesni +16
Highly capable AI systems could secretly pursue misaligned goals -- what we call "scheming". Because a scheming AI would deliberately try to hide its misaligned goals and actions,…
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38
AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
Charlotte Stix, Matteo Pistillo, Girish Sastry +6
The most advanced future AI systems will first be deployed inside the frontier AI companies developing them. According to these companies and independent experts, AI systems may re…
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Tomek Korbak, Mikita Balesni, Buck Shlegeris +1
As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from caus…
Frontier Models are Capable of In-context Scheming
Alexander Meinke, Bronson Schoen, Jérémy Scheurer +3
Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabiliti…
Towards evaluations-based safety cases for AI scheming
Mikita Balesni, Marius Hobbhahn, David Lindner +13
We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through sc…