1 paper · 1 filter
Zixuan Chen, Weikai Lu, Xin Lin +1
Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLM…