4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2024
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
Fabio Pernisi, Dirk Hovy, Paul Röttger
As diverse linguistic communities and users adopt large language models (LLMs), assessing their safety across languages becomes critical. Despite ongoing efforts to make LLMs safe,…
cs.CL2024★ 4 cited
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
Paul Röttger, Fabio Pernisi, Bertie Vidgen +1
The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abun…