works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CY2026

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

Eleanor M. Marshall, Pedro Medeiros, Peter Peneder +6

The paper introduces BioTIER, a benchmark of 542 expert‑curated prompts designed to help large language models refuse high‑risk biological information while still providing benign…

cs.AI2026

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

Andrew Bo Liu, Samira Nedungadi, Bryce Cai +3

Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM…

cs.AI2026

LLM Novice Uplift on Dual-Use, In Silico Biology Tasks

Chen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus +16

Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than…

cs.CR2025

Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models

Boyi Wei, Zora Che, Nathaniel Li +10

Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad acto…

cs.CY2025

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

Tegan McCaslin, Jide Alaga, Samira Nedungadi +5

Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducte…

cs.CY2025

Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Jasper Götting, Pedro Medeiros, Jon G Sanders +6

We present the Virology Capabilities Test (VCT), a large language model (LLM) benchmark that measures the capability to troubleshoot complex virology laboratory protocols. Construc…