1 paper · 1 filter
Xirui Li, Ruochen Wang, Minhao Cheng +2
The safety alignment of Large Language Models (LLMs) is vulnerable to both manual and automated jailbreak attacks, which adversarially trigger LLMs to output harmful content. Howev…