3 papers
cs.CY2025
What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework
Tom Reed, Tegan McCaslin, Luca Righetti
Most frontier AI developers publicly document their safety evaluations of new AI models in model reports, including testing for chemical and biological (ChemBio) misuse risks. This…
cs.CY2025
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
Tegan McCaslin, Jide Alaga, Samira Nedungadi +5
Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducte…
cs.CY2024
AI Safety Frameworks Should Include Procedures for Model Access Decisions
Edward Kembery, Tom Reed
The downstream use cases, benefits, and risks of AI models depend significantly on what sort of access is provided to the model, and who it is provided to. Though existing safety f…