1 citations · 2 across the 14 of their papers we have counts for
1 paper · 2 filters
Yanghan Wang, Zhiqiang Kou, Fu Feng +2
Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vulnerable to malicious prompts…