2 citations · 2 across the 10 of their papers we have counts for
1 paper · 1 filter
Miles Q. Li, Benjamin C. M. Fung, Boyang Li +2
Existing white-box jailbreak attacks against aligned LLMs typically append discrete adversarial suffixes to the user prompt, which visibly alters the prompt and operates in a combi…