From the 1 of 3 linked papers with an AI index.
3 papers
BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation
Eleanor M. Marshall, Pedro Medeiros, Peter Peneder +6
The paper introduces BioTIER, a benchmark of 542 expert‑curated prompts designed to help large language models refuse high‑risk biological information while still providing benign…
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
Chen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus +16
Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than…
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
Jasper Götting, Pedro Medeiros, Jon G Sanders +6
We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Construc…