17 citations · 34 across the 19 of their papers we have counts for
1 paper · 1 filter
Sicheng Zhu, Ruiyi Zhang, Bang An +6
Safety alignment of Large Language Models (LLMs) can be compromised with manual jailbreak attacks and (automatic) adversarial attacks. Recent studies suggest that defending against…