papers

Publications (6)

cs.CL2024

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97

This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…

cs.CY2024

To Err is AI : A Case Study Informing LLM Flaw Reporting Practices

Sean McGregor, Allyson Ettinger, Nick Judd +10

In August of 2024, 495 hackers generated evaluations in an open-ended bug bounty targeting the Open Language Model (OLMo) from The Allen Institute for AI. A vendor panel staffed by…

cs.SI2022

Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation

Stefan Wojcik, Sophie Hilgard, Nick Judd +5

We present an approach for selecting objectively informative and subjectively helpful annotations to social media posts. We draw on data from on an online environment where contrib…

cs.CR2025

SandboxEval: Towards Securing Test Environment for Untrusted Code

Rafiqul Rabin, Jesse Hostetler, Sean McGregor +2

While large language models (LLMs) are powerful assistants in programming tasks, they may also produce malicious code. Testing LLM-generated code therefore poses significant risks…

cs.CR2025

Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation

Rafiqul Rabin, Sean McGregor, Nick Judd

This paper explores the risk that a large language model (LLM) trained for code generation on data mined from software repositories will generate content that discloses sensitive i…

cs.HC2025

Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk

Nick Judd, Alexandre Vaz, Kevin Paeth +5

We introduce findings and methods to facilitate evidence-based discussion about how large language models (LLMs) should behave in response to user signals of risk of suicidal thoug…