3 papers
cs.LG2025
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset
Gary Ackerman, Theodore Wilson, Zachary Kallenborn +7
The potential for rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons…
cs.LG2025
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models II: Benchmark Generation Process
Gary Ackerman, Zachary Kallenborn, Anna Wetzel +7
The potential for rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LLMs), to facilitate bioterrorism or access to biological weapons…
cs.LG2025
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
Gary Ackerman, Brandon Behlendorf, Zachary Kallenborn +7
Both model developers and policymakers seek to quantify and mitigate the risk of rapidly-evolving frontier artificial intelligence (AI) models, especially large language models (LL…