1 paper · 1 filter
Kristina NikoliÄ, Luze Sun, Jie Zhang +1
Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are act…