1 paper · 1 filter
Yangyang Guo, Ziwei Xu, Si Liu +2
This study reveals a previously unexplored vulnerability in the safety alignment of Large Language Models (LLMs). Existing aligned LLMs predominantly respond to unsafe queries with…