From the 1 of 6 linked papers with an AI index.
6 papers
BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation
Eleanor M. Marshall, Pedro Medeiros, Peter Peneder +6
The paper introduces BioTIER, a benchmark of 542 expert‑curated prompts designed to help large language models refuse high‑risk biological information while still providing benign…
ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity
Andrew Bo Liu, Samira Nedungadi, Bryce Cai +3
Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM…
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
Chen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus +16
Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than…
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
Boyi Wei, Zora Che, Nathaniel Li +10
Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad acto…
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
Tegan McCaslin, Jide Alaga, Samira Nedungadi +5
Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducte…
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
Jasper Götting, Pedro Medeiros, Jon G Sanders +6
We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Construc…