collaborators

7 papers

cs.LG2025

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Hjalmar Wijk, Tao Lin, Joel Becker +20

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations fo…

cs.LG2025

An Example Safety Case for Safeguards Against Misuse

Joshua Clymer, Jonah Weinbaum, Robert Kirk +3

Existing evaluations of AI misuse safeguards provide a patchwork of evidence that is often difficult to connect to real-world decisions. To bridge this gap, we describe an end-to-e…

cs.CY2025

Bare Minimum Mitigations for Autonomous AI Development

Joshua Clymer, Isabella Duan, Chris Cundy +10

Artificial intelligence (AI) is advancing rapidly, with the potential for significantly automating AI research and development itself in the near future. In 2024, international sci…

cs.AI2025

A sketch of an AI control safety case

Tomek Korbak, Joshua Clymer, Benjamin Hilton +2

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how devel…

cs.CR2024

Towards evaluations-based safety cases for AI scheming

Mikita Balesni, Marius Hobbhahn, David Lindner +13

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through sc…

cs.CL2024

Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals

Joshua Clymer, Caden Juang, Severin Field

Like a criminal under investigation, Large Language Models (LLMs) might pretend to be aligned while evaluated and misbehave when they have a good opportunity. Can current interpret…