4 papers
Global Cybercrime Damages: A Baseline for Frontier AI Risk Assessment
KamilÄ LukoÅ¡iÅ«tÄ, John Halstead, Luca Righetti
AI companies and governments are increasingly concerned about frontier AI systems enabling cybercrime, yet defining meaningful capability thresholds requires knowing the scale of c…
Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
Shen Zhou Hong, Alex Kleinman, Alyssa Mathiowetz +9
Large language models (LLMs) perform strongly on biological benchmarks, raising concerns that they may help novice actors acquire dual-use laboratory skills. Yet, whether this tran…
What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework
Tom Reed, Tegan McCaslin, Luca Righetti
Most frontier AI developers publicly document their safety evaluations of new AI models in model reports, including testing for chemical and biological (ChemBio) misuse risks. This…
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
Tegan McCaslin, Jide Alaga, Samira Nedungadi +5
Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducte…