collaborators

6 papers

cs.AI2025

AI Deception: Risks, Dynamics, and Controls

Boyuan Chen, Sitong Fang, Jiaming Ji +56

As intelligence increases, so does its shadow. AI deception, in which systems induce false beliefs to secure self-beneficial outcomes, has evolved from a speculative concern to an…

cs.AI2025

The Singapore Consensus on Global AI Safety Research Priorities

Yoshua Bengio, Tegan Maharaj, Luke Ong +84

Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…

cs.CY2025

Promising Topics for U.S.-China Dialogues on AI Risks and Governance

Saad Siddiqui, Lujain Ibrahim, Kristy Loke +5

Cooperation between the United States and China, the world's leading artificial intelligence (AI) powers, is crucial for effective global AI governance and responsible AI developme…

cs.CY2025

Bare Minimum Mitigations for Autonomous AI Development

Joshua Clymer, Isabella Duan, Chris Cundy +10

Artificial intelligence (AI) is advancing rapidly, with the potential for significantly automating AI research and development itself in the near future. In 2024, international sci…

cs.CY2025

In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?

Ben Bucknall, Saad Siddiqui, Lara Thurnherr +19

International cooperation is common in AI research, including between geopolitical rivals. While many experts advocate for greater international cooperation on AI safety to address…

cs.CL2025

Comparative Analysis of Efficient Adapter-Based Fine-Tuning of State-of-the-Art Transformer Models

Saad Mashkoor Siddiqui, Mohammad Ali Sheikh, Muhammad Aleem +1

In this work, we investigate the efficacy of various adapter architectures on supervised binary classification tasks from the SuperGLUE benchmark as well as a supervised multi-clas…