3 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Xuhao Hu, Peng Wang, Xiaoya Lu +3
Previous research has shown that LLMs finetuned on malicious or incorrect completions within narrow domains (e.g., insecure code or incorrect medical advice) can become broadly mis…