5 citations · 5 across the 4 of their papers we have counts for
5 papers
AI Value Alignment for Evolving Social Norms
Nenad Tomašev, Matija Franklin, Simon Osindero
AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop…
From AGI to ASI
Tim Genewein, Matija Franklin, Alexander Lerchner +11
Over the last decade, building human-level artificial general intelligence has moved from far-fetched speculation to being a concrete next-decade target for many of the largest AI…
Distributional AGI Safety
Nenad Tomašev, Matija Franklin, Julian Jacobs +2
AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithi…
Intelligent AI Delegation
Nenad Tomašev, Matija Franklin, Simon Osindero
AI agents are able to tackle increasingly complex tasks. To achieve more ambitious goals, AI agents need to be able to meaningfully decompose problems into manageable sub-component…
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
Stuart Armstrong, Matija Franklin, Connor Stevens +1
Recent work showed Best-of-N (BoN) jailbreaking using repeated use of random augmentations (such as capitalization, punctuation, etc) is effective against all major large language…