1 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Zouying Cao, Yifei Yang, Hai Zhao
Safety alignment is indispensable for Large Language Models (LLMs) to defend threats from malicious instructions. However, recent researches reveal safety-aligned LLMs prone to rej…