8 papers
Role Steering of Language Models for Social Simulations
Isaac Song, Mohammed Rehan Parwani, Glenn Matlin +8
Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activat…
Positive Alignment: Artificial Intelligence for Human Flourishing
Ruben Laukkonen, Seb Krier, Chloé Bakalar +13
Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…
Distributional AGI Safety
Nenad Tomašev, Matija Franklin, Julian Jacobs +2
AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithi…
Comprehensive AI governance requires addressing non-model gains
Arthur Goemans, Dan Altman, Noemi Dreksler +8
Frontier AI governance often centres on the model-level governance paradigm, which assumes that a model's capability profile is primarily a function of the compute and data used du…
Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt
Joel Z. Leibo, Alexander Sasha Vezhnevets, William A. Cunningham +3
Artificial Intelligence (AI) systems are increasingly placed in positions where their decisions have real consequences, e.g., moderating online spaces, conducting research, and adv…
An Approach to Technical AGI Safety and Security
Rohin Shah, Alex Irpan, Alexander Matt Turner +27
Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…