Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
Nanna Inie, Jonathan Stray, Leon Derczynski
Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition o…
cs.CL2024
Introducing v0.5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97
This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…