6 citations · 11 across the 25 of their papers we have counts for
1 paper · 1 filter
Zhuohang Long, Siyuan Wang, Shujun Liu +3
Jailbreak attacks, where harmful prompts bypass generative models' built-in safety, raise serious concerns about model vulnerability. While many defense methods have been proposed,…