889 citations · 1.7k across the 41 of their papers we have counts for
1 paper · 1 filter
Raffaele Mura, Giorgio Piras, Kamilė Lukošiūtė +3
Jailbreaks are adversarial attacks designed to bypass the built-in safety mechanisms of large language models. Automated jailbreaks typically optimize an adversarial suffix or adap…