9 citations · 9 across the 1 of their papers we have counts for
2 papers
cs.AI2024★ 44 cited
The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence
Peter Slattery, Alexander K. Saeri, Emily A. C. Grundy +7
Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet researchers, policymakers, and technology c…
cs.CL2023★ 9 cited
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Rusheb Shah, Quentin Feuillade--Montixi, Soroush Pour +3
Despite efforts to align large language models to produce harmless responses, they are still vulnerable to jailbreak prompts that elicit unrestricted behaviour. In this work, we in…