2 citations · 4 across the 9 of their papers we have counts for
4 papers · 1 filter
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38
AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…
An alignment safety case sketch based on debate
Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton +1
If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feed…
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Tomek Korbak, Mikita Balesni, Buck Shlegeris +1
As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from caus…
A sketch of an AI control safety case
Tomek Korbak, Joshua Clymer, Benjamin Hilton +2
As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how devel…