1 paper · 1 filter
Yanghan Wang, Zhiqiang Kou, Fu Feng +2
Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vulnerable to malicious prompts…