5.3k citations · 10.5k across the 32 of their papers we have counts for
7 papers · 1 filter
International AI Safety Report 2026
Yoshua Bengio, Stephen Clare, Carina Prunkl +89
The International AI Safety Report 2026 synthesises the current scientific evidence on the capabilities, emerging risks, and safety of general-purpose AI systems. The report series…
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies
Miles Brundage, Noemi Dreksler, Aidan Homewood +45
We outline a vision for frontier AI auditing, which we define as rigorous third-party verification of frontier AI developers' safety and security claims, and evaluation of their sy…
AI Behind Closed Doors: a Primer on The Governance of Internal Deployment
Charlotte Stix, Matteo Pistillo, Girish Sastry +6
The most advanced future AI systems will first be deployed inside the frontier AI companies developing them. According to these companies and independent experts, AI systems may re…
International AI Safety Report
Yoshua Bengio, Sören Mindermann, Daniel Privitera +93
The first International AI Safety Report comprehensively synthesizes the current evidence on the capabilities, risks, and safety of advanced AI systems. The report was mandated by…
Computing Power and the Governance of Artificial Intelligence
Girish Sastry, Lennart Heim, Haydn Belfield +16
Computing power, or "compute," is crucial for the development and deployment of artificial intelligence (AI) capabilities. As a result, governments and companies have started to le…
Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations
Josh A. Goldstein, Girish Sastry, Micah Musser +3
Generative language models have improved drastically, and can now produce realistic text outputs that are difficult to distinguish from human-written content. For malicious actors,…