1 paper · 1 filter
Xiangman Li, Xiaodong Wu, Qi Li +2
Jailbreak attacks pose a serious threat to the safety of Large Language Models (LLMs) by crafting adversarial prompts that bypass alignment mechanisms, causing the models to produc…