1 paper · 1 filter
Haoming Wen, Shi Chen, Qingyu Shi +4
Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with only a few steps of supervised…