1 paper · 1 filter
Yuping Lin, Pengfei He, Han Xu +4
Large language models (LLMs) are susceptible to a type of attack known as jailbreaking, which misleads LLMs to output harmful contents. Although there are diverse jailbreak attack…