1 citations · 2 across the 7 of their papers we have counts for
1 paper · 1 filter
Raffaele Mura, Giorgio Piras, KamilÄ LukoÅ¡iÅ«tÄ +3
Jailbreaks are adversarial attacks designed to bypass the built-in safety mechanisms of large language models. Automated jailbreaks typically optimize an adversarial suffix or adap…