1 citations · 3 across the 5 of their papers we have counts for
9 papers · 1 filter
International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management
Yoshua Bengio, Stephen Clare, Carina Prunkl +66
This second update to the 2025 International AI Safety Report assesses new developments in general-purpose AI risk management over the past year. It examines how researchers, publi…
Agentic Misalignment: How LLMs Could Be Insider Threats
Aengus Lynch, Benjamin Wright, Caleb Larson +5
We stress-tested 16 leading models from multiple developers in hypothetical corporate environments to identify potentially risky agentic behaviors before they cause real harm. In t…
International AI Safety Report 2025: First Key Update: Capabilities and Risk Implications
Yoshua Bengio, Stephen Clare, Carina Prunkl +70
Since the publication of the first International AI Safety Report, AI capabilities have continued to improve across key domains. New training techniques that teach AI systems to re…
The Singapore Consensus on Global AI Safety Research Priorities
Yoshua Bengio, Tegan Maharaj, Luke Ong +84
Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…
Bare Minimum Mitigations for Autonomous AI Development
Joshua Clymer, Isabella Duan, Chris Cundy +10
Artificial intelligence (AI) is advancing rapidly, with the potential for significantly automating AI research and development itself in the near future. In 2024, international sci…
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
Ben Bucknall, Saad Siddiqui, Lara Thurnherr +19
International cooperation is common in AI research, including between geopolitical rivals. While many experts advocate for greater international cooperation on AI safety to address…