4 citations · 8 across the 5 of their papers we have counts for
6 papers
The Future of the AI Summit Series
Lucia Velasco, Charles Martinet, Henry de Zoete +27
This policy memo examines the evolution of the international AI Summit series, initiated at Bletchley Park in 2023 and continued through Seoul in 2024 and Paris in 2025, as a forum…
International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management
Yoshua Bengio, Stephen Clare, Carina Prunkl +66
This second update to the 2025 International AI Safety Report assesses new developments in general-purpose AI risk management over the past year. It examines how researchers, publi…
SpecEval: Evaluating Model Adherence to Behavior Specifications
Ahmed Ahmed, Kevin Klyman, Yi Zeng +2
Companies that develop foundation models publish behavioral guidelines they pledge their models will follow, but it remains unclear if models actually do so. While providers such a…
International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
Yoshua Bengio, Stephen Clare, Carina Prunkl +70
Since the publication of the first International AI Safety Report, AI capabilities have continued to improve across key domains. New training techniques that teach AI systems to re…
The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis
Eoin O'Doherty, Nicole Weinrauch, Andrew Talone +4
Artificial intelligence (AI) is advancing at a pace that raises urgent questions about how to align machine decision-making with human moral values. This working paper investigates…
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
Elise Karinshak, Amanda Hu, Kewen Kong +4
Immense effort has been dedicated to minimizing the presence of harmful or biased generative content and better aligning AI output to human intention; however, research investigati…