11 papers
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
Haoming Wen, Shi Chen, Qingyu Shi +4
Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised…
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel +185
Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…
AICrypto: Evaluating Cryptography Capabilities of Large Language Models
Yu Wang, Yijian Liu, Liheng Ji +11
We build \textbf{AICrypto}, a comprehensive benchmark designed to evaluate the cryptography capabilities of large language models (LLMs). The benchmark comprises 135 multiple-choic…
Reverse-Engineering Model Editing on Language Models
Zhiyu Sun, Minrui Luo, Yu Wang +2
Large language models (LLMs) are pretrained on corpora containing trillions of tokens and, therefore, inevitably memorize sensitive information. Locate-then-edit methods, as a main…
Can Large Language Models Reinvent Foundational Algorithms?
Jian Zhao, Haoren Luo, Yu Wang +3
LLMs have shown strong potential to advance scientific discovery. Whether they possess the capacity for foundational innovation, however, remains an open question. In this work, we…
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
Baicheng Chen, Yu Wang, Ziheng Zhou +4
Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly prone to vulnerabilities. It supports cr…