6 papers
Automated alignment is harder than you think
Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau +1
A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue th…
An alignment safety case sketch based on debate
Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton +1
If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feed…
Emerging Practices in Frontier AI Safety Frameworks
Marie Davidsen Buhl, Ben Bucknall, Tammy Masterson
As part of the Frontier AI Safety Commitments agreed to at the 2024 AI Seoul Summit, many AI developers agreed to publish a safety framework outlining how they will manage potentia…
Safety Cases: A Scalable Approach to Frontier AI Safety
Benjamin Hilton, Marie Davidsen Buhl, Tomek Korbak +1
Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g.…
Safety case template for frontier AI: A cyber inability argument
Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett +4
Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offerin…
Safety cases for frontier AI
Marie Davidsen Buhl, Gaurav Sett, Leonie Koessler +2
As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do s…