Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
An alignment safety case sketch based on debate
Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton +1
If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feed…
cs.AI2025
A sketch of an AI control safety case
Tomek Korbak, Joshua Clymer, Benjamin Hilton +2
As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how devel…