1 paper
Yitong Guo, Xiaoyi Chen, Siyuan Zhang +2
Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this fai…