6 papers
AI Deception: Risks, Dynamics, and Controls
Boyuan Chen, Sitong Fang, Jiaming Ji +56
As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an…
The Singapore Consensus on Global AI Safety Research Priorities
Yoshua Bengio, Tegan Maharaj, Luke Ong +84
Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…
Promising Topics for U.S.-China Dialogues on AI Risks and Governance
Saad Siddiqui, Lujain Ibrahim, Kristy Loke +5
Cooperation between the United States and China, the world's leading artificial intelligence (AI) powers, is crucial for effective global AI governance and responsible AI developme…
Bare Minimum Mitigations for Autonomous AI Development
Joshua Clymer, Isabella Duan, Chris Cundy +10
Artificial intelligence (AI) is advancing rapidly, with the potential for significantly automating AI research and development itself in the near future. In 2024, international sci…
In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
Ben Bucknall, Saad Siddiqui, Lara Thurnherr +19
International cooperation is common in AI research, including between geopolitical rivals. While many experts advocate for greater international cooperation on AI safety to address…
Comparative Analysis of Efficient Adapter-Based Fine-Tuning of State-of-the-Art Transformer Models
Saad Mashkoor Siddiqui, Mohammad Ali Sheikh, Muhammad Aleem +1
In this work, we investigate the efficacy of various adapter architectures on supervised binary classification tasks from the SuperGLUE benchmark as well as a supervised multi-clas…