1 paper · 1 filter
Xuhao Hu, Peng Wang, Xiaoya Lu +3
Previous research has shown that LLMs finetuned on malicious or incorrect completions within narrow domains (e.g., insecure code or incorrect medical advice) can become broadly mis…