3 citations · 3 across the 1 of their papers we have counts for
1 paper · 1 filter
Canaan Yung, Hanxun Huang, Christopher Leckie +1
Adversarial prompts are capable of jailbreaking frontier large language models (LLMs) and inducing undesirable behaviours, posing a significant obstacle to their safe deployment. C…