4 papers
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
Jianwei Li, Jung-Eun Kim
Backdoor attacks pose severe security threats to large language models (LLMs), where a model behaves normally under benign inputs but produces malicious outputs when a hidden trigg…
Superficial Safety Alignment Hypothesis
Jianwei Li, Jung-Eun Kim
As large language models (LLMs) are overwhelmingly more and more integrated into various applications, ensuring they generate safe responses is a pressing need. Previous studies on…
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
Xingli Fang, Jianwei Li, Varun Mulchandani +1
The capabilities of artificial intelligence systems have been advancing to a great extent, but these systems still struggle with failure modes, vulnerabilities, and biases. In this…
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
Jianwei Li, Jung-Eun Kim
Recent studies on the safety alignment of large language models (LLMs) have revealed that existing approaches often operate superficially, leaving models vulnerable to various adve…