activity
20242026
collaborators

6 papers

cs.AI2026

Automated alignment is harder than you think

Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau +1

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue th…

cs.AI2025

An alignment safety case sketch based on debate

Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton +1

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feed…

cs.CY2025

Emerging Practices in Frontier AI Safety Frameworks

Marie Davidsen Buhl, Ben Bucknall, Tammy Masterson

As part of the Frontier AI Safety Commitments agreed to at the 2024 AI Seoul Summit, many AI developers agreed to publish a safety framework outlining how they will manage potentia…

cs.CY2025

Safety Cases: A Scalable Approach to Frontier AI Safety

Benjamin Hilton, Marie Davidsen Buhl, Tomek Korbak +1

Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g.…

cs.CY2024

Safety case template for frontier AI: A cyber inability argument

Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett +4

Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offerin…

cs.CY2024

Safety cases for frontier AI

Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler +2

As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do s…