30 citations · 43 across the 10 of their papers we have counts for
1 paper · 1 filter
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu +5
The safety alignment of current Large Language Models (LLMs) is vulnerable. Relatively simple attacks, or even benign fine-tuning, can jailbreak aligned models. We argue that many…