10 citations · 10 across the 3 of their papers we have counts for
5 papers
Foundation models may exhibit staged progression in novel CBRN threat disclosure
Kevin M Esvelt
The extent to which foundation models can disclose novel chemical, biological, radiation, and nuclear (CBRN) threats to expert users is unclear due to a lack of test cases. I lever…
Responsible Reporting for Frontier AI Development
Noam Kolt, Markus Anderljung, Joslyn Barnhart +7
Mitigating the risks from frontier AI systems requires up-to-date and reliable information about those systems. Organizations that develop and deploy frontier systems have signific…
A system capable of verifiably and privately screening global DNA synthesis
Carsten Baum, Jens Berlips, Walther Chen +28
Printing custom DNA sequences is essential to scientific and biomedical research, but the technology can be used to manufacture plagues as well as cures. Just as ink printers recog…
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal +54
The White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and che…
Will releasing the weights of future large language models grant widespread access to pandemic agents?
Anjali Gopal, Nathan Helm-Burger, Lennart Justen +6
Large language models can benefit research and human understanding by providing tutorials that draw on expertise from many different fields. A properly safeguarded model will refus…