Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
How to evaluate control measures for LLM agents? A trajectory from today to superintelligence
Tomek Korbak, Mikita Balesni, Buck Shlegeris +1
As LLM agents grow more capable of causing harm autonomously, AI developers will rely on increasingly sophisticated control measures to prevent possibly misaligned agents from caus…
cs.AI2025
A sketch of an AI control safety case
Tomek Korbak, Joshua Clymer, Benjamin Hilton +2
As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how devel…