activity
20242026
collaborators

9 papers

cs.LG2026

Predicting LLM Safety Before Release by Simulating Deployment

Marcus Williams, Hannah Sheahan, Cameron Raymond +8

Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model beha…

cs.LG2025

Async Control: Stress-testing Asynchronous Control Measures for LLM Agents

Asa Cooper Stickland, Jan Michelfeit, Arathi Mani +6

LLM-based software engineering agents are increasingly used in real-world development tasks, often with access to sensitive data or security-critical codebases. Such agents could i…

cs.CL2025

Lessons from Studying Two-Hop Latent Reasoning

Mikita Balesni, Tomek Korbak, Owain Evans

Large language models can use chain-of-thought (CoT) to externalize reasoning, potentially enabling oversight of capable LLM agents. Prior work has shown that models struggle at tw…

cs.AI2025

How to evaluate control measures for LLM agents? A trajectory from today to superintelligence

Tomek Korbak, Mikita Balesni, Buck Shlegeris +1

As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from caus…

cs.CY2025

Safety Cases: A Scalable Approach to Frontier AI Safety

Benjamin Hilton, Marie Davidsen Buhl, Tomek Korbak +1

Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g.…

cs.AI2025

A sketch of an AI control safety case

Tomek Korbak, Joshua Clymer, Benjamin Hilton +2

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how devel…