2 citations · 2 across the 3 of their papers we have counts for
3 papers
An alignment safety case sketch based on debate
Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton +1
If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feed…
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Tomek Korbak, Mikita Balesni, Buck Shlegeris +1
As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from caus…
Safety case template for frontier AI: A cyber inability argument
Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett +4
Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offerin…