4 papers
The AI Risk Spectrum: From Dangerous Capabilities to Existential Threats
Markov Grey, Charbel-Raphaël Segerie
As AI systems become more capable, integrated, and widespread, understanding the associated risks becomes increasingly important. This paper maps the full spectrum of AI risks, fro…
The bitter lesson of misuse detection
Hadrien Mariaccia, Charbel-Raphaël Segerie, Diego Dorn
Prior work on jailbreak detection has established the importance of adversarial robustness for LLMs but has largely focused on the model ability to resist adversarial inputs and to…
Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
Markov Grey, Charbel-Raphaël Segerie
As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform govern…
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
Ben Bucknall, Saad Siddiqui, Lara Thurnherr +19
International cooperation is common in AI research, including between geopolitical rivals. While many experts advocate for greater international cooperation on AI safety to address…