collaborators

8 papers

cs.CL2026

Role Steering of Language Models for Social Simulations

Isaac Song, Mohammed Rehan Parwani, Glenn Matlin +8

Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activat…

cs.AI2026

Positive Alignment: Artificial Intelligence for Human Flourishing

Ruben Laukkonen, Seb Krier, Chloé Bakalar +13

Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psych…

cs.AI2026

Distributional AGI Safety

Nenad Tomašev, Matija Franklin, Julian Jacobs +2

AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithi…

cs.CY2026

Comprehensive AI governance requires addressing non-model gains

Arthur Goemans, Dan Altman, Noemi Dreksler +8

Frontier AI governance often centres on the model-level governance paradigm, which assumes that a model's capability profile is primarily a function of the compute and data used du…

cs.AI2025

Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt

Joel Z. Leibo, Alexander Sasha Vezhnevets, William A. Cunningham +3

Artificial Intelligence (AI) systems are increasingly placed in positions where their decisions have real consequences, e.g., moderating online spaces, conducting research, and adv…

cs.AI2025

An Approach to Technical AGI Safety and Security

Rohin Shah, Alex Irpan, Alexander Matt Turner +27

Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…