9 papers
Predicting LLM Safety Before Release by Simulating Deployment
Marcus Williams, Hannah Sheahan, Cameron Raymond +8
Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model beha…
Async Control: Stress-testing Asynchronous Control Measures for LLM Agents
Asa Cooper Stickland, Jan Michelfeit, Arathi Mani +6
LLM-based software engineering agents are increasingly used in real-world development tasks, often with access to sensitive data or security-critical codebases. Such agents could i…
Lessons from Studying Two-Hop Latent Reasoning
Mikita Balesni, Tomek Korbak, Owain Evans
Large language models can use chain-of-thought (CoT) to externalize reasoning, potentially enabling oversight of capable LLM agents. Prior work has shown that models struggle at tw…
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Tomek Korbak, Mikita Balesni, Buck Shlegeris +1
As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from caus…
Safety Cases: A Scalable Approach to Frontier AI Safety
Benjamin Hilton, Marie Davidsen Buhl, Tomek Korbak +1
Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g.…
A sketch of an AI control safety case
Tomek Korbak, Joshua Clymer, Benjamin Hilton +2
As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how devel…